Ramp Scales GPU AI Workloads Using AWS ECS Managed Instances
Ramp, a finance automation platform, details its transition to Amazon ECS Managed Instances for scaling GPU-powered AI inference workloads. The shift from manually managed EC2 fleets significantly reduced operational overhead and unified their GPU infrastructure with existing ECS setups. This allows Ramp to deliver reliable, real-time AI at scale for tasks like fraud prevention and transaction categorization. The post highlights how ECS Managed Instances combines EC2 flexibility with Fargate's simplicity, managing instance provisioning, scaling, and patching for GPU instances like NVIDIA L40S and A10G.
- →Ramp's initial manual EC2 management for GPU workloads
- →Adopting Amazon ECS Managed Instances for GPU workloads
- →Ramp's Bore internal ML inference platform architecture
- →Multi-model GPU inference in Ramp's Embeddings service
- →Consistent cluster and capacity provider layout for GPU services
Notes (5) ›
- Ramp's initial manual EC2 management for GPU workloads
Ramp initially used an "ECS on EC2" pattern for GPU workloads, involving manual provisioning of Auto Scaling groups, launch templates, AMIs, and monitoring scripts. This approach led to significant operational overhead, inconsistent patching, and separate management from their main ECS infrastructure.
- Adopting Amazon ECS Managed Instances for GPU workloads
Ramp transitioned to Amazon ECS Managed Instances, a fully managed compute option combining EC2 flexibility with Fargate's simplicity. AWS now handles instance configuration, scaling, patching, and maintenance, integrating GPU infrastructure seamlessly with their existing ECS environment.
- Ramp's Bore internal ML inference platform architecture
Bore, Ramp’s internal ML inference platform, uses a mixed compute pattern: CPU-bound services run on Fargate, while GPU-backed services like text similarity and merchant matching run on ECS Managed Instances, targeting NVIDIA L40S GPUs using g6e instances in production.
- Multi-model GPU inference in Ramp's Embeddings service
The Embeddings service hosts multiple vector embedding models, provisioning dedicated capacity providers for different hardware requirements. It uses NVIDIA A10G (g5.xlarge) for default models and NVIDIA L40S (g6e instances) for larger models like Qwen3-Embedding and Gemma-300M.
- Consistent cluster and capacity provider layout for GPU services
Ramp organizes its ECS infrastructure with a base Terraform workspace per environment, provisioning cluster-level resources and capacity providers. GPU services declare a capacity provider strategy pointing to specific providers, while non-GPU services continue to use Fargate within the same clusters.
https://aws.amazon.com/blogs/containers/how-ramp-runs-gpu-ai-workloads-at-scale-with-ecs-managed-instances/
Related releases
- Amazon EVS Achieves FedRAMP Class C Compliance in US Regions AWS What's New ·
- Amazon EC2 I7ie Instances Now Available in AWS Israel (Tel Aviv) Region AWS What's New ·
- Amazon EC2 X8i Instances Now Available in AWS São Paulo Region AWS What's New ·
- AWS introduces Amazon EC2 T8i instances, offering improved price performance AWS What's New ·
- AWS launches EC2 T8i instances for low-cost burstable workloads AWS News Blog ·
- Amazon WorkSpaces adds support for NVIDIA Blackwell GPU instances AWS What's New ·