Gemma-4-31B LLMs from Google DeepMind and NVIDIA now on SageMaker JumpStart
Google DeepMind's Gemma-4-31B-it-assistant and NVIDIA's Gemma-4-31B-IT-NVFP4 models are now available on Amazon SageMaker JumpStart. These models bring the flagship Gemma 4 31B dense architecture to enterprise workloads in both full-precision and optimized quantized variants. The assistant-tuned model offers multimodal reasoning, coding, and agentic workflows with a 256K-token context window. The NVIDIA-optimized version delivers up to 2.5x faster inference and 68% reduced memory footprint, making it ideal for cost-efficient, high-throughput production deployments.
Features (1) ›
- Gemma-4-31B-IT-NVFP4 model with NVIDIA optimization
This model delivers the same Gemma 4 31B capabilities, optimized with NVIDIA's ModelOpt framework to 4-bit FP4 precision. It reduces memory usage by 68% to ~18.5 GB and achieves approximately 2.5x faster inference while retaining 97–99% of the original model's quality.
https://aws.amazon.com/about-aws/whats-new/2026/01/gemma-4-31b-it-assistant-gemma-4-31b-it-nvfp4-jumpstart/
Related releases
- SageMaker JumpStart adds NVIDIA Qwen3.6-35B-A3B-NVFP4 and Alibaba Wan2.1-T2V-1.3B models AWS What's New ·
- Amazon SageMaker JumpStart Adds Granite-Speech, Kanana-2, and OpenFold3 Models AWS What's New ·
- Mistral AI's Ministral-3 Models Now on Amazon SageMaker JumpStart AWS What's New ·
- SageMaker HyperPod adds model caching for faster inference autoscaling AWS What's New ·
- Amazon SageMaker Feature Store now supports individual feature updates AWS What's New ·
- Amazon SageMaker Batch Transform adds support for G6e instances AWS What's New ·