Understanding the AI That Drives Robots
An Explanation of Vision-Language-Action Models
Over a year has passed since I last examined the progress of humanoid robots, and excitement for this technology remains high. Startups NEURA Robotics, Figure AI, Apptronik, Unitree, and Agility Robotics have collectively raised over $2 billion in funding. China’s Unitree went public, raising approximately $900 million, while Agility Robotics plans to go public via SPAC later this year.
The World Humanoid Robot Games, held in China, showcased robots performing various physical tasks. However, the impressive demos we have seen must be viewed with caution.
One notable demonstration was Figure's robots sorting packages and autonomously unloading a dishwasher. Physical Intelligence showcased its robot model making coffee and folding boxes in a chocolate factory. Generalist AI demonstrated its GEN-1.5 model learning new tasks from just a few examples.
Advancements in hardware, such as improved actuators, have contributed to some progress in robotics. However, the majority of advancements come from improvements in robot AI, which utilizes specially developed AI models to control robots. As more general AI continues to rapidly advance, it is likely that we will see a similar acceleration in robotic AI capabilities. Currently, robotic AI seems quite limited, but given the rapid advancement of general AI, this could change quickly.
Understanding how the AI models used to control robots function is essential. There are several types of robotic AIs, often referred to as "policies," which map a particular set of inputs to a set of robot actions. For this essay, we will focus on one commonly used architecture called a vision-language-action model (VLA), employed by companies such as Figure, Unitree, Physical Intelligence, and Nvidia.
Specifically, we will examine Physical Intelligence's open-weight π0.5 VLA, which was released in 2025 and has gained widespread use (although it is not the company's most advanced model).
A VLA takes text, images, and robot information as input and produces a series of robot actions as output. It employs many of the same components that made large language models (LLMs) successful, primarily attention and the transformer architecture. While it is not clear if VLAs will remain the primary robotic AI paradigm, they are currently a widely used method for controlling robots.
Written by urgent.news from Construction Physics's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.