aws AWS What's New ·

Amazon SageMaker HyperPod Inference Gateway for Scalable LLM Inference

aiawsgaengineeraws-eksaws-sagemaker
feature

AWS has launched the SageMaker HyperPod Inference Gateway, a Kubernetes-native, GPU-aware routing system designed for scalable LLM inference. Deployed as an EKS managed add-on, it replaces traditional load balancing with real-time inference-signal-driven routing, significantly reducing first-token latency and p99 Time-To-First-Token (TTFT) by up to 82% and 98% respectively. The gateway works with any OpenAI-compatible model server, enabling a single endpoint to serve multiple models without client-side changes. It is now generally available in all AWS Regions that support the SageMaker HyperPod inference add-on.

Features (1) ›
  • SageMaker HyperPod Inference Gateway for LLM Scalability

    The Amazon SageMaker HyperPod Inference Gateway is a new Kubernetes-native, GPU-aware routing system for large language model (LLM) inference. It deploys as an EKS managed add-on, providing real-time, inference-signal-driven routing that substantially reduces first-token latency and improves p99 TTFT in mixed-hardware and burst traffic scenarios. The gateway supports any OpenAI-compatible model server and allows a single endpoint to serve multiple models with no application code changes.

Read the original announcement →

https://aws.amazon.com/about-aws/whats-new/2026/09/sagemaker-hyperpod-inference-gateway/

Related releases