aws AWS Containers Blog ·

Ramp Scales GPU AI Workloads Using AWS ECS Managed Instances

blogaiawsarchitectaws-ec2aws-ecs
announcement

Ramp, a finance automation platform, details its transition to Amazon ECS Managed Instances for scaling GPU-powered AI inference workloads. The shift from manually managed EC2 fleets significantly reduced operational overhead and unified their GPU infrastructure with existing ECS setups. This allows Ramp to deliver reliable, real-time AI at scale for tasks like fraud prevention and transaction categorization. The post highlights how ECS Managed Instances combines EC2 flexibility with Fargate's simplicity, managing instance provisioning, scaling, and patching for GPU instances like NVIDIA L40S and A10G.

  • Ramp's initial manual EC2 management for GPU workloads
  • Adopting Amazon ECS Managed Instances for GPU workloads
  • Ramp's Bore internal ML inference platform architecture
  • Multi-model GPU inference in Ramp's Embeddings service
  • Consistent cluster and capacity provider layout for GPU services
Notes (5)
  • Ramp's initial manual EC2 management for GPU workloads

    Ramp initially used an "ECS on EC2" pattern for GPU workloads, involving manual provisioning of Auto Scaling groups, launch templates, AMIs, and monitoring scripts. This approach led to significant operational overhead, inconsistent patching, and separate management from their main ECS infrastructure.

  • Adopting Amazon ECS Managed Instances for GPU workloads

    Ramp transitioned to Amazon ECS Managed Instances, a fully managed compute option combining EC2 flexibility with Fargate's simplicity. AWS now handles instance configuration, scaling, patching, and maintenance, integrating GPU infrastructure seamlessly with their existing ECS environment.

  • Ramp's Bore internal ML inference platform architecture

    Bore, Ramp’s internal ML inference platform, uses a mixed compute pattern: CPU-bound services run on Fargate, while GPU-backed services like text similarity and merchant matching run on ECS Managed Instances, targeting NVIDIA L40S GPUs using g6e instances in production.

  • Multi-model GPU inference in Ramp's Embeddings service

    The Embeddings service hosts multiple vector embedding models, provisioning dedicated capacity providers for different hardware requirements. It uses NVIDIA A10G (g5.xlarge) for default models and NVIDIA L40S (g6e instances) for larger models like Qwen3-Embedding and Gemma-300M.

  • Consistent cluster and capacity provider layout for GPU services

    Ramp organizes its ECS infrastructure with a base Terraform workspace per environment, provisioning cluster-level resources and capacity providers. GPU services declare a capacity provider strategy pointing to specific providers, while non-GPU services continue to use Fargate within the same clusters.

Read the original announcement →

https://aws.amazon.com/blogs/containers/how-ramp-runs-gpu-ai-workloads-at-scale-with-ecs-managed-instances/

Related releases