Introducing Gemini Robotics ER 2
Gemini Robotics ER 2 is a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
Google has unveiled Gemini Robotics ER 2, a cutting-edge model intended to serve as a sophisticated "brain" for robots. This advanced system enables real-time spatial reasoning, multi-step task planning, and coordination between various robots. Developers can now access Gemini Robotics ER 2 through the Gemini API, Google AI Studio, or the Gemini Enterprise Agent Platform, empowering them to create their own physical AI agents.
The model is designed to facilitate robots' ability to comprehend their surroundings, make rapid decisions, and collaborate with other robots. It features the capability to analyze video feeds to monitor their own performance and rectify errors in real-time. Moreover, it stands out for its emphasis on enhancing robots' safety and practicality in the physical world.
While high-resolution spatial reasoning is crucial, ER 2 goes a step further by enabling robots to think and react swiftly, in harmony with the speed of the physical world. This capability is a key factor for robots aiming to assist humans in everyday settings.
ER 2 stands as a substantial leap from its predecessor, ER 1.6. It empowers robots to anticipate future actions while simultaneously executing them. A notable improvement is the model's capacity to monitor continuous video feeds, enabling robots to assess their progress, respond to unforeseen situations, and smoothly transition to subsequent steps.
Gemini Robotics ER 2 is now publicly accessible to developers via the Gemini API, Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform. To assist users, Google has furnished examples on configuring the model and prompting it to execute more efficient physical AI tasks.
In complex tasks requiring multiple steps, ER 2 acts as a physical agent, orchestrating the robot's actions and enabling self-correction and generalization to novel situations. Developers can define low-level control interfaces, such as Vision-Language-Action (VLA) models or navigation APIs, as tools within the system, and stream multimodal video, audio, or text directly into the model.
To assess ER 2's performance, Google has conducted evaluations using both simulated and real-world robot control, as well as pairing the model with human remote control. In every scenario, ER 2 demonstrated superior performance compared to its predecessor, ER 1.6.
This advanced model integrates seamlessly with the Gemini Live API, a bidirectional streaming endpoint optimized for latency-sensitive tasks. The outcome is seamless orchestration, where ER 2 commands action models and robotics APIs to accomplish multi-step tasks without interruptions. A demonstration of this capability is showcased with Spot from Boston Dynamics, which fetched a popcorn snack based on natural language commands, courtesy of ER 2.
ER 2 addresses one of robotics' most significant challenges: determining when a task is complete. The model introduces significant advancements in video understanding and progress tracking. It quantifies task progress by classifying video frames into five progress levels (0-20%, 20-40%, 40-60%, 60-80%, 80-100%). This real-time situational awareness empowers robots to adjust actions on the fly or retry failed steps without restarting an entire workflow.
Progress classification accuracy for ER 2 is 57.4%, surpassing previous models and competing frontier models. Additionally, the model excels in moment-finding, accurately identifying critical video frames—such as the perfect moment to stop pouring coffee into a cup—with 91.3% accuracy and a mean absolute distance of just 0.96 seconds. This level of precision offers a substantial performance boost compared to larger model categories, all while maintaining sub-second latency crucial for safe real-world robotics operation.
Written by urgent.news from Google Blog's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.