Gemini Robotics 2 Brings Google's AI Into the Physical World
The latest version of Google DeepMind's AI model includes a significant jump into “physical AGI.” But plopping AI into the real world comes with risks.
Gemini Robotics 2 unifies multiple AI models into a single system, enabling robots to comprehend their environment and act accordingly. A vision language model (VLM) interprets images and video, communicating with humans and strategizing tasks. Two vision language action (VLA) models, trained for physical movement and robotic manipulation, guide a robot’s full-body movement and the actions of grippers or hands.
Google DeepMind trained the amalgamated model using human teleoperation, video examples, and simulations, noting that AI models require specific training to perform diverse complex tasks. While Anthropic and OpenAI lead in chatbots and AI coding tools, Google's robotics research is more robust, with past collaborations involving Boston Dynamics.
Gemini Robotics 2 signifies Google's belief that AI must transcend the digital realm to reach its full potential, aiming for "physical AGI" – robots capable of anything humans can do. However, integrating frontier AI models into robots poses risks, as previous research reveals unexpected and potentially dangerous behaviors. Google addresses safety concerns through a multi-layered approach, including guardrails on each model layer and ASIMOV-Agentic, a new benchmark to measure the safety of AI systems collaborating to control a robot.
CEO Demis Hassabis envisions developing an AI operating system for various robots, akin to Android for smartphones.
Written by urgent.news from Wired Business's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.