Benchmark gives 4 models control of a Corolla, 1 finishes the course
Three researchers say they let language models steer a real car around a cone course, and published the token bill for each attempt.
A benchmark published at drivingbench.com hands language models direct control of a Toyota Corolla's steering and pedals on a fixed cone course. Of the four models on its leaderboard, only GPT-6 Astra finished. The paper has not verified the runs, and no one else appears to have reproduced them.
The site credits Aditya Ramabadran, Simon Mahns and Tobias Gessler with equal contributions, and describes the work as experimental software with no official affiliation to any car or AI company. Each model gets up to three attempts in one continuous chat, according to the page.
GPT-6 Astra completed the course on its second attempt in 5 minutes 22 seconds, the leaderboard says, after reaching 49% on the first. It puts the best progress of Claude Fable 5.1 at 45%, Grok 4.6 at 11% and GPT-5.6 Sol at 6%.
The finishing run consumed 246.6 million tokens and cost $7.74, according to the same table. Scoring counts how far a model travels along the course centreline while staying within 4 metres of it. Completion time, distance covered, token use and cost also feature.
The page does not say when the runs took place, what prompt or scaffolding each model got, how often it answered, or whether anyone sat in the car. The submission had 67 points and 41 comments on the Hacker News front page an hour after it appeared.