Google Cloud Boosts AI Infrastructure with Filestore, GKE, and New Security Tools
Google Cloud announced several August and July updates to its AI infrastructure and orchestration offerings. Key enhancements include a new Filestore backend on Colossus for independent IOPS provisioning and deep GKE integration, as well as dedicated Cloud Run instances for cost-effective, always-on AI agents. The updates also bring general availability for Managed Lustre and C4N VMs, significantly scaling GKE Dataplane V2 to 15,000 nodes with network policies, and introducing an open-source tool, k8s-aibom, for AI supply chain security. These changes aim to improve performance, scalability, and cost efficiency for demanding AI and machine learning workloads on Google Cloud.
- →gVisor sandboxes available for distributed Ray clusters on GKE
- →Dedicated Cloud Run instances for cost-effective AI agents
- →Google Cloud Managed Lustre generally available
- →GKE Dataplane V2 scales up to 15,000 nodes with Network Policies
- →Open-source k8s-aibom tool for AI supply chain security on GKE
Security (1) ›
- Open-source k8s-aibom tool for AI supply chain security on GKE
Google Cloud open-sourced k8s-aibom, a lightweight Kubernetes controller that monitors clusters to detect running AI runtimes and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). This tool helps secure AI supply chains, deploy AI workloads safely, and mitigate shadow AI.
Features (4) ›
- gVisor sandboxes available for distributed Ray clusters on GKE
Google Cloud introduced an experimental library in partnership with Anyscale, bringing gVisor, an open-source application kernel, directly into distributed Ray clusters on GKE. This provides lightweight environments with stronger isolation than ordinary containers, alongside fast startup times and low memory overhead.
- Dedicated Cloud Run instances for cost-effective AI agents
New Cloud Run instances offer dedicated, singleton compute runtimes that remain active even when idle, designed for personal AI agents. These instances are offered at a low cost, making them an economical option for continuous AI workloads.
- Google Cloud Managed Lustre generally available
Google Cloud Managed Lustre is now generally available in four performance tiers, providing throughput from 125 MB/s to 1000 MB/s per TiB, and scaling up to 8 PB of capacity. This solution, powered by DDN’s EXAScaler, offers high-performance storage for demanding AI workloads.
- GKE Dataplane V2 scales up to 15,000 nodes with Network Policies
GKE Dataplane V2 now supports standard GKE clusters scaling up to 15,000 nodes, maintaining full active Network Policy enforcement. This capability addresses the infrastructure needs of large enterprise and AI/ML customers.
Enhancements (1) ›
- Filestore now optimized for AI and agentic workflows
Filestore now leverages a new backend built directly on Colossus, enabling independent provisioning of IOPS from storage capacity. This enhancement, deeply integrated with GKE, is designed to support agentic swarms requiring high-performance access to common datasets.
https://cloud.google.com/blog/topics/ai-infrastructure/whats-new-in-ai-infrastructure-this-month/
Related releases
- Google details BigQuery Graph's GA and cross-cloud analytics for AI agents Google Cloud Blog ·
- Google Cloud details agentic data pipelines for MLOps with the Data Agent Kit Google Cloud Blog ·
- Looker Update Brings AI Data Agents, In-Database Analytic Models, and CI Enhancements Google Cloud release notes ·
- BigQuery Enhances Compliance, Graph Processing, and XGBoost ML Capabilities Google Cloud release notes ·
- BigQuery Addresses Security Vulnerability and Adds Real-time Python UDF Logging Google Cloud release notes ·
- Terraform Google Provider v8.0.0 introduces significant breaking changes Terraform Google Provider Releases ·