Gemini Robotics ER 2 Enhances Embodied Reasoning for Robots
Google has launched Gemini Robotics ER 2, an updated embodied reasoning model for robots designed to improve real-time decision-making and task planning in physical environments. This release enhances robots' ability to interact with humans, understand their surroundings, and execute multi-step tasks with improved temporal intelligence and adaptability. It is now available to developers via the Gemini API and Google AI Studio, with broader access on Gemini Enterprise Agent Platform.
- →Gemini Robotics ER 2: Advanced Embodied Reasoning Model
- →Multi-Robot Collaboration Capabilities
- →Improved Task Execution and Adaptability
- →Advancements in Temporal Intelligence
- →Enhanced Spatial and Safety Intelligence
Features (2) ›
- Gemini Robotics ER 2: Advanced Embodied Reasoning Model
Gemini Robotics ER 2 is introduced as Google's most capable embodied reasoning model, designed to enable robots to think and act in real-time, bridging high-level reasoning with motor execution. The model allows robots to chat with humans, understand the physical world, and plan multi-step tasks by handing off execution to vision-language-action models and natively calling tools.
- Multi-Robot Collaboration Capabilities
The new release introduces multi-robot collaboration, allowing diverse robot types to communicate and work together in shared spaces. This enables the completion of complex workflows that would be beyond the capabilities of a single robot.
Enhancements (3) ›
- Improved Task Execution and Adaptability
Gemini Robotics ER 2 offers significant upgrades over its predecessor by enabling robots to track their progress via continuous video feeds, adapt to errors, and know when to proceed to the next task. This allows for more fluid orchestration without jarring pauses, even during complex, multi-step processes.
- Advancements in Temporal Intelligence
Gemini Robotics ER 2 significantly improves temporal understanding for robust task completion, featuring progress classification with 57.4% accuracy and moment-finding at 91.3% accuracy. This allows robots to precisely identify task completion points and manage sequential actions effectively.
- Enhanced Spatial and Safety Intelligence
The model advances spatial reasoning with improved success/failure detection on raw video, broader instrument reading capabilities, and enhanced visual question answering. It also demonstrates significant gains in safety, halting operations when humans are near and adhering to physical constraints during reasoning tasks.
Notes (1) ›
- Availability and Developer Resources
Gemini Robotics ER 2 is publicly available to developers via the Gemini API and Google AI Studio, with a private preview on the Gemini Enterprise Agent Platform. Code examples and configuration guides are provided to facilitate adoption.
https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/
Related releases
- Vibe-coded schedule app built with Gemini simplifies mornings Gemini Blog ·
- Gemini App Updates: July Drop Gemini Blog ·
- Gemini Spark integrates Chrome for automated web errands Gemini Blog ·
- Python GenAI SDK v2.16.0 adds environment resource, Maps tool support Gemini Python SDK Releases ·
- Google Python GenAI SDK Adds Audio Transcription and Enterprise Mode Gemini Python SDK Releases ·
- Gemini for macOS adds voice-based natural language features Gemini Blog ·