Google's GKE & Cloud Run gain AI optimizations, new security features; cited as Gartner Leader
Google has introduced significant enhancements to GKE and Cloud Run, focusing on AI and agentic application support. These updates improve performance and efficiency for AI inference workloads, boost startup times, and introduce new security features like Agent Substrate and Agent Sandbox for untrusted code execution. Key additions include serverless GPU support on Cloud Run, GKE autoscaling based on application intent, and specialized Filestore volumes for agents, targeting engineers and architects building complex AI systems. The company's efforts are validated by its recognition as a Leader in the 2026 Gartner Magic Quadrant for Container Management.
- →Enhanced security and scalability for Kubernetes-based agentic infrastructure
- →New developer experience capabilities for serverless containers with Cloud Run
- →Google recognized as a Leader in 2026 Gartner Magic Quadrant for Container Management
- →Leading performance and efficiency for AI infrastructure on GKE and Cloud Run
Features (2) ›
- Enhanced security and scalability for Kubernetes-based agentic infrastructure
New features include the open-source GKE Agent Substrate for secure agent execution, and GKE Agent Sandbox built on gVisor for isolating untrusted AI code. GKE Dataplane V2 scalability limits doubled to 15,000 nodes, and GKE gains intent-based autoscaling, reacting faster to application needs, along with Filestore agent volumes for rapid NFS mounts.
- New developer experience capabilities for serverless containers with Cloud Run
Cloud Run now offers one-click prototyping directly within Google AI Studio and introduces 'Cloud Run instances' for cost-effective, long-running singleton resources with integrated Cloud Storage mounts. Additionally, Cloud Run sandboxes provide hard-isolated environments for safely executing model-generated code with fast spin-up times.
Enhancements (1) ›
- Leading performance and efficiency for AI infrastructure on GKE and Cloud Run
GKE introduces a predictive latency boost and automatic KV Cache storage tiering to reduce Time-to-First-Token by up to 70% and 40% respectively. Container and model startup times on GKE are up to 4x faster, and Cloud Run now supports on-demand NVIDIA RTX PRO 6000 Blackwell GPUs with scale-to-zero capabilities for large models.
Notes (1) ›
- Google recognized as a Leader in 2026 Gartner Magic Quadrant for Container Management
Google was positioned highest in Ability to Execute, ranking first in every use case in the accompanying Critical Capabilities report for container management. This marks Google's fourth consecutive year as a Leader in the category.
https://cloud.google.com/blog/products/containers-kubernetes/2026-gartner-magic-quadrant-for-container-management/
Related releases
- Google Cloud launches managed service for Gemini RL fine-tuning Google Cloud Blog ·
- Latin American Midsize Businesses Drive Digital Transformation with Google Cloud AI Google Cloud Blog ·
- Google Cloud details why AI startups choose its comprehensive stack Google Cloud Blog ·
- Google Cloud Networking Supports Fluid Compute Choices for AI Workloads Google Cloud Blog ·
- Google Cloud SDK 586.0.0 Includes Breaking Changes and New Features Google Cloud release notes ·
- Lucius AI runs global tender platform with AlloyDB and MCP, achieves 47x search speedup Google Cloud Blog ·