i
News
News · 2026-10-07

OpenAI’s Astra drove a Toyota Corolla to In-N-Out

@neuronium_ai @neuronium_ai

A team of engineers put OpenAI’s Astra behind the wheel of a Toyota Corolla and sent it to an In-N-Out. The trip was a side project, not a self-driving product test: the model was built to generate text, and the engineers had not prepared it for the drive. Nothing alarming happened. But letting a general-purpose model control a moving car is a striking test of whether AI can reason about the physical world—and a reminder that the ability to act beyond a screen carries risks as well as promise.

Cover: OpenAI’s Astra drove a Toyota Corolla to In-N-Out

The physical-world gap

AI models can answer difficult questions and handle some tasks in virtual environments. Outside computers and the internet, they still struggle to make sense of their surroundings. Even as companies talk about artificial general intelligence, physical reasoning remains an open problem.

Some researchers have left large companies to work on it. Andrew Dai, head of Elorian AI and a former Google DeepMind researcher, argues that better visual reasoning could enable systems to judge whether restaurant guests like their food, or help robots work in homes. Robotics, he says, is a crucial test: a useful home robot would need to understand the physical world.

Elorian and Scale AI recently created Humanity’s Sixth Sense, a benchmark for measuring how well models understand physical scenes. Scale AI research scientist Xingang Guo, who helped build it, said the team wanted to move beyond perception and test whether a model could intuitively understand a scene as people do.

That is a different bar from recognizing objects. A model must use what it sees to decide what to do next—and revise that decision when the world does not behave as expected.

From a parking lot to a burger run

The driving experiment began at weekend meetups. The engineers noticed that models such as Astra could generate complex 3D simulations and wondered whether that ability might help them navigate reality. They also live in the San Francisco Bay Area, where Teslas and Waymos are common.

The group—Ramabadhran, Mans and Gessler—first tested Grok from SpaceXAI, then moved on to newer models from OpenAI and Anthropic. At first, the models refused to control the cars, saying they could interpret road images but could not issue commands to a real vehicle. Carefully chosen prompts got them to continue.

The team says major AI companies are focusing on spatial reasoning in newer models, but it doubts the models were specifically trained to drive. Mans’s guess is that driving ability may have emerged from scaling multimodal training on images, video and 3D models.

Astra was the only model to complete the simple route in the team’s DrivingBench test, which takes place on a marked-out parking-lot course. It did so very slowly. Claude Fable 5.1 covered 45 percent of the route; Grok managed 11 percent.

100percent route
45Claude Fable
11Grok

The engineers say the models adapted as they drove, correcting mistakes and improving their control in context. That is notable, but the benchmark results also put the burger run in perspective: completing one trip is not the same as passing a driving test.

What the drive does—and does not—show

The most interesting claim here is not that a language model can drive a car. It is that skills developed for reasoning across images and 3D scenes may transfer to a physical task without explicit driving preparation. I think that makes the experiment useful as a probe, not evidence that these models are ready for the road.

The announcement is quiet about how reliably that adaptation works beyond this trip, or what happens when a model makes a mistake it cannot correct. The engineers acknowledge the risk of giving a general-purpose model control of a fast-moving car. As models gain more ability to act in the physical world, the distance between an impressive demonstration and a safe system matters more than the demonstration itself.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X