aws AWS Big Data Blog ·

Spark Connect now supported on Amazon EMR on EC2 for interactive PySpark

blogdataawsgaengineerhealthcareaws-ec2aws-sagemaker
feature announcement

Amazon EMR on EC2 now supports Spark Connect, enabling interactive PySpark development and debugging from local IDEs or SageMaker Unified Studio Data Notebooks. This feature separates the application code from the Spark engine, allowing developers to debug against full-size data directly on an EMR cluster, which resolves prior challenges with environment mismatches and slow code iteration. It allows teams to share a single EMR cluster for up to 1,000 concurrent, isolated sessions, improving resource utilization and consistency. This capability requires Amazon EMR with emr-spark-8.0 or later and Python 3.9+ with `pyspark[connect]` locally.

  • →Interactive PySpark development with Spark Connect on EMR on EC2
  • →Client-server architecture separates local code from Spark engine
  • →Share a single EMR cluster for team-wide interactive PySpark sessions
  • →Prerequisites and steps to get started with Spark Connect
Features (2) ›
  • Interactive PySpark development with Spark Connect on EMR on EC2

    AWS has announced support for Spark Connect on Amazon EMR on EC2 (emr-spark-8.0+), enabling interactive PySpark development and debugging from local IDEs or SageMaker Unified Studio Data Notebooks. This allows developers to set breakpoints and inspect DataFrames against full-size data directly on a dedicated EMR cluster.

  • Client-server architecture separates local code from Spark engine

    Spark Connect employs a client-server architecture where a lightweight PySpark client runs locally and sends DataFrame/SQL operations via gRPC/TLS to a Spark Connect Server on the EMR cluster. This ensures that the debugging environment matches the production runtime, eliminating environment mismatches and improving the developer experience.

Enhancements (1) ›
  • Share a single EMR cluster for team-wide interactive PySpark sessions

    Teams can now share one dedicated Amazon EMR cluster for up to 1,000 concurrent, isolated Spark Connect sessions, each with its own execution role and tags. This increases resource utilization, ensures consistent environments, and allows for per-user cost attribution and data access isolation.

Notes (1) ›
  • Prerequisites and steps to get started with Spark Connect

    To use Spark Connect, users need an Amazon EMR cluster running emr-spark-8.0.0 or later with SessionEnabled set to true, Python 3.9+ with `pyspark[connect]` installed locally, and specific IAM permissions. The setup involves creating a session-enabled cluster, starting a session, and connecting from an IDE or SageMaker Unified Studio.

Read the original announcement →

https://aws.amazon.com/blogs/big-data/announcing-spark-connect-on-amazon-emr-on-ec2-interactive-pyspark-anywhere/

Related releases