UiPath Scales Agentic AI with Shared GPU Fleet on Google Cloud
UiPath has re-architected its infrastructure to handle the computational demands of agentic AI by implementing a shared GPU fleet on Google Cloud. This move addresses challenges with spiky workloads, supply bottlenecks for high-end GPUs, and operational overhead. The new architecture utilizes a mix of A3 and G4 VM instances managed by Google Kubernetes Engine to balance training and inference needs, enabling predictable costs and faster model deployment for enterprise customers.
- →Google Cloud AI Hypercomputer Facilitates Scalable AI
- →UiPath's Agentic AI Push Requires Advanced Infrastructure
- →Shared GPU Fleet Optimizes Compute Resource Utilization
- →Dynamic Workload Scheduler Secures GPU Capacity
- →Hybrid VM Instance Strategy Balances Cost and Performance
Features (1) ›
- Google Cloud AI Hypercomputer Facilitates Scalable AI
UiPath leveraged the Google Cloud AI Hypercomputer, an integrated system of performance-optimized hardware and flexible software, to support its growing AI scale. This environment minimizes friction between hardware and software, allowing engineering teams to focus on model performance.
Enhancements (3) ›
- Shared GPU Fleet Optimizes Compute Resource Utilization
UiPath transitioned from isolated GPU clusters to a shared Google Cloud GPU fleet, managed by its machine learning services platform. This approach prioritizes work across teams and time windows, balancing inference during the day with batch training at night to maximize utilization and reduce contention.
- Dynamic Workload Scheduler Secures GPU Capacity
To overcome supply bottlenecks for high-end GPUs, UiPath uses Google Cloud's Dynamic Workload Scheduler (DWS) to schedule training runs in advance and secure capacity. This allows for predictable capacity planning rather than reacting to scarcity, ensuring consistent access to necessary compute resources.
- Hybrid VM Instance Strategy Balances Cost and Performance
UiPath now utilizes both A3 VM instances with NVIDIA H100 GPUs for heavy-duty training and fine-tuning, alongside G4 VM instances with NVIDIA RTX Pro 6000 for cost-effective inference workloads. This mix optimizes for both speed and cost, especially for lighter inference tasks.
Notes (2) ›
- UiPath's Agentic AI Push Requires Advanced Infrastructure
UiPath is pioneering agentic AI, requiring significant computational power and reliable infrastructure to deploy autonomous agents for complex business processes. This shift necessitates efficient orchestration of hundreds of GPUs for both massive training jobs and real-time inference without escalating costs or latency.
- Improved AI Capabilities Drive Business Outcomes
With consistent access to Google Cloud GPUs, UiPath can bring advanced models into production, enabling high-accuracy data extraction from unstructured documents. Customers like Omega Healthcare and Thermo Fisher Scientific are experiencing significant improvements in processing time and accuracy.
https://cloud.google.com/blog/topics/customers/how-uipath-built-its-high-performance-gpu-platform/
Related releases
- Google Kubernetes Engine Updates Available Versions and Auto-Upgrade Targets Google Cloud release notes ·
- Compute Engine flexible CUDs are now GA for G2 and G4 GPU accelerator-optimized machine series Google Cloud release notes ·
- Google Distributed Cloud for Bare Metal 1.35.400-gke.81 Now Available Google Cloud release notes ·
- Cloud SDK 580.0.0 Updates Include Deprecation, New Features, and GA Promotions Google Cloud release notes ·
- Google Distributed Cloud (software only) for VMware 1.35.400-gke.81 is now available Google Cloud release notes ·
- Google Cloud outlines PQC roadmap, targets full quantum readiness by 2029 Google Cloud Blog ·