aws AWS What's New ·

Gemma-4-31B LLMs from Google DeepMind and NVIDIA now on SageMaker JumpStart

aiawsgaengineeraws-sagemaker
feature

Google DeepMind's Gemma-4-31B-it-assistant and NVIDIA's Gemma-4-31B-IT-NVFP4 models are now available on Amazon SageMaker JumpStart. These models bring the flagship Gemma 4 31B dense architecture to enterprise workloads in both full-precision and optimized quantized variants. The assistant-tuned model offers multimodal reasoning, coding, and agentic workflows with a 256K-token context window. The NVIDIA-optimized version delivers up to 2.5x faster inference and 68% reduced memory footprint, making it ideal for cost-efficient, high-throughput production deployments.

Features (1)
  • Gemma-4-31B-IT-NVFP4 model with NVIDIA optimization

    This model delivers the same Gemma 4 31B capabilities, optimized with NVIDIA's ModelOpt framework to 4-bit FP4 precision. It reduces memory usage by 68% to ~18.5 GB and achieves approximately 2.5x faster inference while retaining 97–99% of the original model's quality.

Read the original announcement →

https://aws.amazon.com/about-aws/whats-new/2026/01/gemma-4-31b-it-assistant-gemma-4-31b-it-nvfp4-jumpstart/

Related releases