Gemini Robotics 2 brings whole body intelligence to robots
Google has unveiled Gemini Robotics ER 2, a groundbreaking model that serves as a sophisticated "brain" for robots. This advanced technology empowers robots with real-time spatial reasoning, multi-step task planning, and collaboration among various robots. The model can be accessed through the Gemini API, Google AI Studio, or the Gemini Enterprise Agent Platform, enabling users to develop their own physical AI agents.
Unlike its predecessor, ER 2 introduces video understanding capabilities, allowing robots to track their progress and rectify mistakes in real time. This model is specifically designed to enhance robots' safety and effectiveness in real-world environments. By integrating high-level reasoning with swift execution, ER 2 enables robots to effectively assist humans in everyday scenarios.
Gemini Robotics ER 2 serves as a high-level brain for robots, allowing them to communicate with humans, interpret the physical world, and plan complex tasks. It then delegates motor execution to lower-level vision-language-action models. Additionally, ER 2 can utilize tools like Google Search or user-defined functions to gather information when needed.
A significant upgrade from ER 1.6, ER 2 boasts continuous video feed analysis, enabling robots to monitor their own advancement, adapt to errors, and determine optimal moments to transition to subsequent steps. The model also facilitates multi-robot collaboration, enabling robots to collaborate in shared spaces and accomplish intricate workflows that a single robot could not accomplish independently.
Developers can now leverage Gemini Robotics ER 2 via the Gemini API, Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform. The company provides examples of configuring the model and prompting it to execute more practical physical AI tasks. Most real-world tasks involve multiple steps, and ER 2 addresses this by orchestrating these steps for robots and enabling them to self-correct and adapt to novel situations.
To assess ER 2's performance, Google conducts evaluations using robots in simulation, real-world robotics control, and remote human control. The model consistently surpasses ER 1.6 for tool orchestration across three control modes: real VLA, simulated VLA, and human tele-op.
Operating at the speed of the physical world is crucial for successful robotics. ER 2 integrates seamlessly into the Gemini Live API, a bidirectional streaming endpoint optimized for latency-sensitive tasks. This integration results in fluid orchestration, where ER 2 directs action models and robotics APIs to execute multi-step tasks without the disruptive pauses typically associated with robot decision-making.
In a demonstration using Boston Dynamics' Spot robot, ER 2 orchestrated Spot's APIs, including navigation and manipulator movements. This collaboration allowed Spot to fetch a popcorn snack upon receiving a natural language command. The code for this demonstration is available on GitHub, along with other examples.
One of the most challenging aspects of robotics is determining when a task is complete. ER 2 addresses this issue by introducing significant improvements in video understanding and progress tracking. It incorporates two foundational capabilities: progress classification and moment finding. Progress classification enables robots to track their advancement towards task completion by categorizing video frames into five progress levels.
This feature provides robots with real-time situational awareness, allowing them to adjust actions on the fly or retry failed steps without restarting a workflow entirely.
Moment finding measures a model's ability to pinpoint the exact video frame where a critical event occurs. ER 2 achieves impressive accuracy in this domain, reaching 91.3% accuracy in moment finding tasks and achieving a 0.96-second mean absolute distance. ER 2 outperforms larger model categories while delivering this precision at a fraction of the compute cost and 4x the execution speed, providing the sub-second latency required for safe, real-world robotic operations.
Gemini Robotics ER 2 enables multi-robot collaboration, allowing diverse machines to communicate via a shared semantic understanding to delegate tasks and accomplish complex workflows. Apptronik's Apollo 2 and Franka F3 Duo showcase how ER 2 facilitates successful collaboration between different robot models.
In conclusion, Google's Gemini Robotics ER 2 represents a remarkable leap forward in robotic technology. By combining real-time spatial reasoning, multi-step task planning, video understanding, and multi-robot collaboration, ER 2 empowers robots to operate more efficiently, safely, and effectively in various environments. With its vast array of capabilities and practical applications, ER 2 is poised to revolutionize the field of robotics and facilitate the widespread adoption of intelligent, autonomous agents in our daily lives.
Written by urgent.news from Google DeepMind's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.