Source: technologyreview.com
What current systems can do
Google DeepMind’s ALOHA 2 is a modest-looking test bed: two arms, grippers and two cameras. Researchers use it to test Gemini Robotics, an AI system that can perform tasks it has learned from examples, such as packing a simple lunch.
That is real progress. Three years ago, a robot could not do this task. But Gemini Robotics can still fail when asked to do something outside its training data. Edward Jones, a robotics professor at Imperial College London, says a truly general system would need to work across the full range of possible actions. Today, Gemini Robotics can do a few things here and a few things there.

Elon Musk predicts that Tesla’s Optimus humanoid could be one the world’s best-selling products. For now, it’s most often seen handing out food and drinks at Tesla events.
Source: technologyreview.com
The shift behind that progress is from hand-coded instructions to AI systems that interpret a scene and plan movements. Visual-language-action models, or VLAs, learn from images, video and demonstrations of people controlling robots remotely. They can identify objects and attempt tasks they have seen before. The hard part is getting them to cope with tasks and conditions they have not.
Google DeepMind’s latest AI models for robots are increasingly dextrous, if rather slow and erratic
Source: technologyreview.com
The shortage is not just data
One answer is to collect more demonstrations. But there is no equivalent of the huge text collections used to train language models, and each way of gathering physical-world examples comes with a cost:
Former Google DeepMind robotics lead Panag Sanketi, who is now working on his own robotics and AI project, favors combining these sources. Agility Robotics co-founder and robotics chief Jonathan Hurst argues that data alone cannot solve the problem. Everyday tasks are full of variation: making coffee means dealing with different kitchens, machines, cups and ingredients. A VLA would need examples of almost everything a robot might do.
Yann LeCun has made a broader objection: methods that work well for language may not suit the continuous, noisy data robots encounter. One proposed alternative is a world model, trained on video, 3D scans and sensor data to predict how objects move, collide, fall and deform. In principle, accurate simulation could make robot training faster, cheaper and safer. Nvidia and Google are working on the technology, while World Labs and AMI Labs have each raised $1 billion. The field is still young, however, and its basic approaches are still taking shape.
A small test from Physical Intelligence suggests why researchers remain interested. The company’s π0.7 model, released in April 2026, used a lightweight world model to generate images of the steps needed for a task. Asked to put a sweet potato in an air fryer, it hesitated, restarted several times and made a reasonable attempt without completing the job.
PI co-founder Sergey Levine said the result was the first convincing sign that the model could make an acceptable attempt at a task for which researchers had not deliberately collected data or trained it. But the training set did contain two teleoperation examples involving an air fryer, in which a person moved its basket. The model’s apparent leap may have depended on those fragments.
Source: technologyreview.com
Source: technologyreview.com
A demo is not a deployment
Public demonstrations often obscure how much work remains. The robot that joined Nvidia CEO Jensen Huang on stage in March 2025 was remotely controlled by a human, whom its creators described as a “puppeteer behind the scenes.”
Source: technologyreview.com
Even autonomous movement remains hard in unfamiliar, cluttered environments. So does handling a broad instruction such as “make dinner,” which requires a robot to inspect a kitchen, choose ingredients and decide what to do next. Google DeepMind’s attempts to have a robot gather everything needed for mushroom risotto have not succeeded.

How do you teach a robot to use a knife? At the startup Physical Intelligence, it begins with designing the right AI architecture, which includes components dedicated to language, vision and motion. This will help it relate commands — “hey robot, chop my vegetables!” — to appropriate actions.
Source: technologyreview.com

Data to train the AI can come from many sources, but one of the most important is human demonstrations. An employee at the startup controls a robot arm through teleoperation, exposing the AI to the task of slicing a zucchini.
Source: technologyreview.com

Researchers train the AI model on hundred of examples of human demonstrations, as well images from the web and first-person video.
Source: technologyreview.com

Once the AI is trained, the team presents it a task it has not seen before, such as chopping this summer squash. When a robot hasn’t seen the exact task before — it may wonder if that’s a yellow zucchini, or an unusual banana — it can mess up. But any failures can be used to help refine the model.
Source: technologyreview.com
Reliability is another bar. Boston Dynamics founder Marc Raibert says a 70% success rate can be a major improvement on 50%, but it is still a failure for practical use.
Agility says hundreds of its robots are being tested at facilities operated by GXO Logistics, Amazon and Schaeffler. For now, they do simple work such as moving containers and boxes. Hurst says it took years to make the robots safe enough for logistics companies to consider using them.
Musk said in May 2025 that thousands of Optimus robots would be working in Tesla factories by the end of that year. In January, he said some were already performing simple factory tasks. At home, the gap is larger still: 1X Neo, priced at $20,000, can be preordered, but most of its actions currently require a remote human operator. Hurst’s estimate for fully autonomous household robots, when pressed to give one, is ten years.
I think the more revealing benchmark is not whether a robot can complete a novel task once, but whether it can do so safely and reliably without a person close by. The announcement is quiet about that rate—and without it, capability demos say little about when these machines will become useful.
Source: technologyreview.com
The hardware is advancing, but the market may not wait for the most capable design. Chinese companies produced nearly 90% of the roughly 15,000 humanoid robots delivered in 2025, according to Omdia and Unitree. Unitree, which delivered more humanoid robots than any other manufacturer last year, sells a model for less than $6,000. AP reported that Chinese robots are being ordered mainly by corporate and university labs and state-owned enterprises.
That points to a tension in the humanoid pitch: a human-shaped robot may fit human spaces, but it also has to justify its cost and complexity. For now, the machines finding real-world trials are doing narrow, controlled jobs—not the open-ended work that makes the bigger forecasts compelling.
The old promise, with better tools
The ambition is centuries old. Leonardo da Vinci sketched a mechanical knight in 1495. Westinghouse’s seven-foot Elektro appeared at the 1939 New York World’s Fair; WABOT-1, built by Waseda University in Japan, followed in 1973. Honda introduced ASIMO in 2000, but shut the project down in 2018 after it failed to move beyond demonstrations into useful work.
These robots were impressive engineering. None was ready to operate independently in the real world. Today’s systems have more powerful AI behind them, but the central problem remains: a robot must act reliably in a physical environment that changes from one moment to the next.
Demonstrations often make robots look useful, but the machines still mess up far too often to be used reliably in homes and factories.
Source: technologyreview.com
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X