Spark Connect now supported on Amazon EMR on EC2 for interactive PySpark
Amazon EMR on EC2 now supports Spark Connect, enabling interactive PySpark development and debugging from local IDEs or SageMaker Unified Studio Data Notebooks. This feature separates the application code from the Spark engine, allowing developers to debug against full-size data directly on an EMR cluster, which resolves prior challenges with environment mismatches and slow code iteration. It allows teams to share a single EMR cluster for up to 1,000 concurrent, isolated sessions, improving resource utilization and consistency. This capability requires Amazon EMR with emr-spark-8.0 or later and Python 3.9+ with `pyspark[connect]` locally.
- →Interactive PySpark development with Spark Connect on EMR on EC2
- →Client-server architecture separates local code from Spark engine
- →Share a single EMR cluster for team-wide interactive PySpark sessions
- →Prerequisites and steps to get started with Spark Connect
Features (2) ›
- Interactive PySpark development with Spark Connect on EMR on EC2
AWS has announced support for Spark Connect on Amazon EMR on EC2 (emr-spark-8.0+), enabling interactive PySpark development and debugging from local IDEs or SageMaker Unified Studio Data Notebooks. This allows developers to set breakpoints and inspect DataFrames against full-size data directly on a dedicated EMR cluster.
- Client-server architecture separates local code from Spark engine
Spark Connect employs a client-server architecture where a lightweight PySpark client runs locally and sends DataFrame/SQL operations via gRPC/TLS to a Spark Connect Server on the EMR cluster. This ensures that the debugging environment matches the production runtime, eliminating environment mismatches and improving the developer experience.
Enhancements (1) ›
- Share a single EMR cluster for team-wide interactive PySpark sessions
Teams can now share one dedicated Amazon EMR cluster for up to 1,000 concurrent, isolated Spark Connect sessions, each with its own execution role and tags. This increases resource utilization, ensures consistent environments, and allows for per-user cost attribution and data access isolation.
Notes (1) ›
- Prerequisites and steps to get started with Spark Connect
To use Spark Connect, users need an Amazon EMR cluster running emr-spark-8.0.0 or later with SessionEnabled set to true, Python 3.9+ with `pyspark[connect]` installed locally, and specific IAM permissions. The setup involves creating a session-enabled cluster, starting a session, and connecting from an IDE or SageMaker Unified Studio.
https://aws.amazon.com/blogs/big-data/announcing-spark-connect-on-amazon-emr-on-ec2-interactive-pyspark-anywhere/
Related releases
- AWS EC2 M8i and M8i-flex Instances Now Available in AWS European Sovereign Cloud AWS What's New ·
- Amazon EC2 C8i and C8i-flex Instances Now Available in AWS European Sovereign Cloud AWS What's New ·
- Amazon EC2 R8i and R8i-flex Instances Now Available in AWS European Sovereign Cloud AWS What's New ·
- Amazon EMR Introduces Long Term Support for Apache Spark with Version 4.1 AWS What's New ·
- Amazon EMR 7.14 is Now Available AWS What's New ·
- Amazon EVS Achieves FedRAMP Class C Compliance in US Regions AWS What's New ·