Urgent.News

What's breaking now, across thousands of outlets.

Culture

Gemini Robotics 2 place le cerveau des humanoïdes chez Google

L'essentiel Google DeepMind a lancé le 30 juillet Gemini Robotics 2, une famille de trois modèles capables de piloter un humanoïde entier, des pieds aux doigts. L'architecture sépare le raisonnement (Gemini Robotics ER 2) de l'exécution motrice, la première couche commandant la seconde comme un outil. ER 2 est ouvert aux développeurs via l'API Gemini et Google AI Studio ; les modèles de contrôle…

Original French Read in English

Abstract editorial illustration

On July 30, Google DeepMind launched Gemini Robotics 2, a family of three models capable of controlling an entire humanoid, from feet to fingers. The architecture separates reasoning (Gemini Robotics ER 2) from motor execution; the first layer commands the second like a tool. ER 2 is open to developers via the Gemini API and Google AI Studio; the motor control models remain reserved for selected partners.

The embedded version adapts to a completely new robot morphology with a few hours of data. When asked to place the mop in the green bucket on the lower shelf, who decides that the robot must walk to the table, bend its knees, then close its fingers at the right spot? Since July 30, Google DeepMind responds with two distinct layers instead of one, a division that will weigh more heavily than the dexterity demonstrations accompanying it.

Gemini Robotics 2 combines a Gemini Robotics ER 2 reasoning cortex and spinal cord, a vision-language-action (VLA) model, and an On-Device 2 variant that runs this conversion locally on the machine itself. The system's neural analogy holds: ER 2 plays the cortex that plans, the VLA plays the spinal cord, and the action. Google DeepMind takes this division of labor: its reasoning layer delegates motor execution to any low-level model, and the developer declares control interfaces and navigation APIs as tools, just as they would Google Search or a homemade function.

Thinking while moving, this two-layer division responds to a very concrete constraint: the physical world doesn't wait. A robot that freezes to think between each gesture becomes useless in a kitchen or workshop. ER 2 therefore relies on the Gemini Live API and its bidirectional streaming entry point, designed to reduce response time.

The model follows a continuous video stream, checks where it is, recovers from a failed gesture, and triggers the next step at the right moment, without the pause of reflection that marked previous generations. Google compares ER 2 to ER 1.6 in tool orchestration across three configurations: a real-action model, a simulated model, and a human remote operator, the latter being the most telling.

The reasoning layer is indifferent to the executor's nature: a VLA, simulation, or human hands all occupy the same spot in its plan. The body has become a parameter. Demonstrations drive the point home. Apollo 2, Apptronik's humanoid, receives a single command and chains walking, flexing, and grasping. The same software layer performs a node, screws a piece, and then coordinates Apollo with a second robot, Duo, to tidy a garage: the reasoning model cuts the task, designates zones, and chooses the relay moment.

At Boston Dynamics, it's Spot that executes, with its navigation and manipulation APIs declared as tools. When a builder arrives with an unprecedented morphology, the embedded version adapts in just a few hours of data. The cost of integrating a new chassis collapses, while value concentrates in the layer that never changes. An open brain, closed muscles confirm this hierarchy.

ER 2's open access to developers via Gemini API and Google AI Studio, in private preview on the Gemini Enterprise Agent Platform, is published on GitHub, while motor control models remain in the hands of carefully selected partners. Competitors' boundaries shift elsewhere: NVIDIA releases weights of its humanoid model GR00T N1.7 under open commercial license, downloadable on Hugging Face, while Figure abandoned OpenAI in early 2025 to develop its own vision-language-action model, Helix.

The consequence is immediate for a robot manufacturer: they can prototype the planning and speaking part, but the muscle-controlling layer doesn't belong to them, and the brain runs at Google. The question of software sovereignty, learned by industrialists for their cloud, resurfaces here in machines that move around in their warehouses.

The success rates shown are carefully chosen. No success rates accompany the tasks demonstrated, and nothing states how many attempts precede the correct mop placement. The published comparative figures focus on tool orchestration compared to the previous version, not the reliability of a humanoid navigating fragile objects and humans.

These reservations don't change the message that the robot industry is now structured like the rest of computing: an operating system on one side, hardware on the other. Humanoid manufacturers who have just learned where they stand in this division of labor have one less edge in the race.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Culture

More German than many Germans

Article URL: https://mertbulan.com/more-german-than-many-germans/ Comments URL: https://news.ycombinator.com/item?id=49151734 Points: 246 # Comments: 161

More from Monday 3 August →