SageMaker HyperPod cuts first-token latency by 82% with Kubernetes-native inference routing
The SageMaker HyperPod Inference Gateway replaces round-robin with real-time signal-driven routing, cutting first-token latency by up to 82% with no application code changes. Deploy it as an EKS add-on if you serve multiple models or a heterogeneous GPU fleet behind a single endpoint.