{
  "id": 72145,
  "title": "Gemini Robotics 2 puts the brain of humanoids at Google",
  "url": "https://urgent.news/2026/08/03/gemini-robotics-2-place-le-cerveau-des-humanoides-chez-google",
  "topic": "culture",
  "section": "Culture",
  "published": "2026-08-03T06:00:08.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/thibault_monteiro/gemini-robotics-2-place-le-cerveau-des-humanoides-chez-google-5k4"
  },
  "original_language": "fr",
  "account": "The essentials Google DeepMind launched Gemini Robotics 2 on July 30, a family of three models capable of controlling an entire humanoid, from feet to fingers. The architecture separates reasoning (Gemini Robotics ER 2) from motor execution, with the first layer commanding the second like a tool. ER 2 is open to developers via the Gemini API and Google AI Studio; motor control models remain reserved for selected partners. The embedded version adapts to a completely new robot morphology with a few hours of data.\n\nYou ask a robot to put the watering can in the green bin, on the bottom shelf. Who decides to walk to the table, bend their knees, and then close their fingers in the right place? Since July 30, Google DeepMind has responded with two distinct layers rather than a single model. This separation will weigh more heavily than the demonstrations of dexterity that accompany it.\n\nA cortex and a spinal cord Gemini Robotics 2 brings together three models. Gemini Robotics ER 2 ensures embodied reasoning: this vision-language model discusses with humans, understands the scene, and breaks down a task into a sequence of actions that can last several minutes. The vision-language-action model (VLA) converts what the robot sees and what it is told into motor commands: walking, crouching, controlling a five-fingered hand. A third variant, Gemini Robotics On-Device 2, performs this conversion locally, on the machine itself.\n\nThe analogy of the nervous system holds quite well. ER 2 plays the cortex that plans, VLA plays the spinal cord and the gesture. Google DeepMind assumes this division of labor: its reasoning layer delegates motor execution to any low-level model. The developer declares control interfaces and navigation APIs as tools, and the brain calls them, exactly as it would call Google Search or a home function.\n\nThinking while moving This two-tiered structure responds to a very concrete constraint: the physical world does not wait. A robot that freezes to think between each gesture becomes unusable in a kitchen or workshop. ER 2 relies on the Gemini Live API and its bidirectional streaming entry point, designed to reduce response time. The model follows a continuous video stream, checks where it is, catches up on a missed gesture, and triggers the next step at the right time, without the reflection pause that betrayed previous generations.\n\nGoogle compares ER 2 to ER 1.6 on tool orchestration in three configurations: a real action model, a simulated model, and a human in teleoperation, i.e., at remote control. The third case is the most striking. The reasoning layer is indifferent to the nature of the executor: a VLA, a simulation, or a pair of human hands occupy the same box in its plan. The body has become a parameter.\n\nHardware, an adjustment variable The demonstrations drive the point home. Apollo 2, Apptronik's humanoid, receives a single instruction and chains walking, bending, and grasping. The same software layer makes a knot, screws a piece, and then cooperates with a second robot, Duo, to organize a garage: the reasoning model breaks down the task, designates areas, and chooses the moment of passing control.\n\nAt Boston Dynamics, it's Spot that executes, with its navigation and manipulation APIs declared as tools. When a manufacturer arrives with a new morphology, the embedded version adapts in a few hours of data. The integration cost of a new chassis collapses, while the value concentrates in the layer that never changes.\n\nThe brain open, the muscles closed The access sharing confirms the hierarchy. ER 2, the brain, is open to developers via the Gemini API and Google AI Studio, in private preview on the Gemini Enterprise Agent Platform, with examples published on GitHub. Motor control models, however, remain in the hands of select partners.\n\nAmong competitors, the boundary lies elsewhere: NVIDIA publishes the weights of its GR00T N1.7 model for humanoids under an open commercial license, downloadable on Hugging Face, while Figure broke with OpenAI in early 2025 to develop its own vision-language-action model, Helix, internally.\n\nThe consequence is immediate for a robot manufacturer: it can prototype the part that plans and speaks, but the layer that commands the muscles does not belong to it, and the brain runs at Google. A humanoid deprived of this layer becomes a teleoperated machine again.\n\nThe question of software sovereignty, which industrialists have learned to ask for their cloud, is replayed here on machines that walk in their warehouses.\n\nSuccess rates are missing The broadcast sequences are carefully chosen. No success rate accompanies the tasks shown, and nothing indicates how many attempts precede the correct taking of the watering can. Published numerical comparisons concern tool orchestration compared to the previous version, not the reliability of a humanoid that moves among fragile objects and humans.\n\nThese reservations do not change what the announcement installs. Robotics is structuring like the rest of IT: an operating system on one side, hardware on the other. Humanoid manufacturers have just learned on which side of this border they are expected.\n\nMy opinion The battle of humanoids will not be won on the joints, and those who only manufacture hardware have already lost it. A manufacturer that connects Gemini Robotics 2 gains two years of development and loses its product: it assembles the terminal of an intelligence updated elsewhere, by someone else. I closely monitor Apptronik, as it is the actor with the strongest...",
  "summary": "L'essentiel Google DeepMind a lancé le 30 juillet Gemini Robotics 2, une famille de trois modèles capables de piloter un humanoïde entier, des pieds aux doigts. L'architecture sépare le raisonnement (Gemini Robotics ER 2) de l'exécution motrice, la première couche commandant la seconde comme un outil. ER 2 est ouvert aux développeurs via l'API Gemini et Google AI Studio ; les modèles de contrôle…",
  "key_points": [
    "Google DeepMind launches Gemini Robotics 2 with ER 2 reasoning layer and VLA model.",
    "ER 2 separates reasoning from motor execution, enabling real-time robot control."
  ],
  "editors_take": null,
  "illustration": "https://urgent.news/ill/72145.png",
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}