aws AWS What's New ·

Amazon SageMaker G7e instances expand to Seoul, London, and Tokyo

aiawsgaengineeraws-ec2aws-sagemaker
feature announcement

Amazon SageMaker AI inference now supports G7e instances in Asia Pacific (Seoul), Europe (London), and Asia Pacific (Tokyo). These instances offer high-performance NVIDIA GPUs and increased memory, crucial for deploying large language models and other generative AI workloads with reduced latency. This expansion enables users to deploy inference endpoints closer to their users in these regions, supporting medium-to-large models efficiently.

  • Amazon SageMaker G7e instances now available in new AWS regions
  • Improved inference performance for generative AI workloads
  • Reduced latency for generative AI workloads
Features (1)
  • Amazon SageMaker G7e instances now available in new AWS regions

    Amazon SageMaker AI inference now supports G7e instances in Asia Pacific (Seoul), Europe (London), and Asia Pacific (Tokyo). These instances feature up to 8 NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs and high-speed networking.

Enhancements (1)
  • Improved inference performance for generative AI workloads

    G7e instances deliver up to 2.3x inference performance compared to G6e instances, with up to 768 GB of total GPU memory. This enables serving large language models up to 70B parameters with FP8 precision without multi-node setups.

Notes (1)
  • Reduced latency for generative AI workloads

    The expansion allows deploying inference endpoints closer to end-users in Asia and Europe, significantly reducing latency for generative AI applications.

Read the original announcement →

https://aws.amazon.com/about-aws/whats-new/2026/07/g7e-sagemaker-ai/

Related releases