A car, a laptop and a safety brake
Three computer scientists built DrivingBench using an internet-connected laptop and a comma four device linked to a Toyota Corolla’s control systems. OpenAI’s GPT-6 Astra, xAI’s Grok 4.6 and Anthropic’s Claude Fable 5.1 received GPS data and vehicle readings, including steering-wheel and wheel angles, then issued steering, acceleration and braking commands.
A human did not drive, but kept a foot above the brake pedal.
The test was a lap around a course in a parking lot. GPT-6 Astra finished in five minutes on its second try. The other models did worse: most attempts ended before the first turn.
The team traced many failures to a basic perception problem. Models misread which side of the first diagonal row of cones marked the driving lane. After its first attempt, Grok said the car was wider than it looked in the camera image: the apparent path ahead was blocked by a flower bed, a wall or the nearest red cone.
A small win with a steep bill
The researchers describe this as their first successful attempt to get frontier models to control a car, following earlier tests with previous generations. Their conclusion is limited: the models can now drive real vehicles at sufficiently low speeds. The result surprised and encouraged the team, while also underscoring the need for more work on safety, alignment and evaluation.
The completed run cost $7.74 for 6.6 million inference tokens. The car was moving at just 0.94 miles per hour and travelled less than 500 feet. The Register calculated that the tokens cost roughly 500 times more than the fuel for a car getting 25 miles per gallon, with gasoline at $4.60 per gallon.
I think the striking part is not that a model completed a course, but how little that accomplishment tells us about driving. The test shows a model can handle one carefully bounded, very slow task; it also shows how easily a misread obstacle can end a run.
The models did not always want to drive
Some models warned the researchers they could not safely control the Toyota. According to the report, some—especially GPT-6 Astra—repeatedly refused on safety grounds, even in an empty parking lot, with a very low speed limit and a human ready to brake.
The team changed its prompts to make the models more willing to take control, at times describing the exercise as a “simulation.” In some trials, however, the models saw real images, recognized that the car was real and became concerned.
That reluctance is part of the result, not a footnote. The models were not specially adapted for driving, and the researchers had to persuade them to act despite their own safety objections. My guess is that the more important next measure is not whether a model can finish a lap, but whether it can reliably recognize when it should not try.
These systems are plainly not ready to commute people to work. Given the safety and legal stakes, that is probably the right outcome for now.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X