Amazon SageMaker HyperPod Inference Gateway for Scalable LLM Inference
AWS has launched the SageMaker HyperPod Inference Gateway, a Kubernetes-native, GPU-aware routing system designed for scalable LLM inference. Deployed as an EKS managed add-on, it replaces traditional load balancing with real-time inference-signal-driven routing, significantly reducing first-token latency and p99 Time-To-First-Token (TTFT) by up to 82% and 98% respectively. The gateway works with any OpenAI-compatible model server, enabling a single endpoint to serve multiple models without client-side changes. It is now generally available in all AWS Regions that support the SageMaker HyperPod inference add-on.
Features (1) ›
- SageMaker HyperPod Inference Gateway for LLM Scalability
The Amazon SageMaker HyperPod Inference Gateway is a new Kubernetes-native, GPU-aware routing system for large language model (LLM) inference. It deploys as an EKS managed add-on, providing real-time, inference-signal-driven routing that substantially reduces first-token latency and improves p99 TTFT in mixed-hardware and burst traffic scenarios. The gateway supports any OpenAI-compatible model server and allows a single endpoint to serve multiple models with no application code changes.
https://aws.amazon.com/about-aws/whats-new/2026/09/sagemaker-hyperpod-inference-gateway/
Related releases
- AWS Named a Leader in Gartner's 2026 Magic Quadrant for Container Management AWS Containers Blog ·
- Standardizing Edge Container Deployments with Amazon EKS: A Strategy Guide AWS Containers Blog ·
- Amazon EMR on EKS now supports Spark Connect for interactive workloads AWS What's New ·
- Amazon EMR on EKS Now Supports IPv6 for Scalable Big Data Workloads AWS What's New ·
- Building a single-pane NOC dashboard for Amazon EKS with Amazon CloudWatch AWS Containers Blog ·
- Amazon EMR Introduces Long Term Support for Apache Spark with Version 4.1 AWS What's New ·