API Gateway supports LLM request routing to Vertex AI models
API Gateway now supports model routing, a managed traffic management layer for LLM requests. This feature allows you to route OpenAI-compatible prompt requests to foundation models in Vertex AI Model Garden, including Gemini, Claude, and GPT models. It centralizes AI traffic management at the network edge and enables in-flight transcoding to standardize client applications. Model routing is configured using new OpenAPI 3.x extensions.
Features (1) ›
- API Gateway Route LLM requests with model routing
Route LLM requests with model routing You can now use model routing in API Gateway as a managed traffic management layer to accept OpenAI-compatible prompt requests, transcode them in-flight, and route them to specific foundation models in Vertex AI Model Garden (including Gemini, Anthropic Claude, and OpenAI GPT models). Key benefits and capabilities include: Centralized traffic management : Consolidate AI traffic routing and lifecycle management at the network edge without hosting standalone client-side proxies. In-flight transcoding : Standardize client applications on an OpenAI-compatible
https://docs.cloud.google.com/release-notes#August_03_2026
Related releases
- Cloud SDK 579.0.0: Breaking Changes, Feature Promotions, and New Commands Google Cloud release notes ·
- Config Connector 1.154.1 Adds New Alpha Resources and Field Support Google Cloud release notes ·
- Google SecOps Marketplace updates Active Directory, Vertex AI, and more Google Cloud release notes ·
- Google Cloud AI and AI Agent Updates: June Recap Google Cloud Blog ·
- Cloud Billing: Early AI workload cost anomaly detection Google Cloud release notes ·
- Vertex AI Search: Lower configurable pricing thresholds Google Cloud release notes ·