Frontier AI Models Test Real-World Driving in Parking Lot Experiment

AI-generated image · US National Wire
Three tech workers used a Toyota Corolla to determine if off-the-shelf LLMs can navigate a physical course without purpose-built training.
A group of San Francisco tech workers known as DrivingBench recently conducted an experiment to see if general-purpose frontier AI models could drive a real vehicle. Unlike systems from Tesla or Waymo, which rely on millions of hours of specific training data, the DrivingBench team used off-the-shelf chatbots running on a laptop to control a rented Toyota Corolla.
According to reporting from 404 Media, the team—consisting of Aditya Ramabadran, Tobias Gessler, and Simon Mahns, who are colleagues at AI math startup Axiom Math—connected the vehicle to the LLMs using Comma, an open-source system that integrates with a car's electrical systems and sensors. The setup used two cameras to feed road data back to the models, which were given control over the brakes, accelerator, and steering.
The team tested four models: GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol. While Grok, Sol, and Fable failed to finish the cone course in a public parking lot, 404 Media reports that GPT-6 Astra eventually learned to navigate and complete the course, though the process required significant troubleshooting.
Ramabadran told 404 Media that the project aimed to create a new benchmark for LLMs performing real-world tasks, noting that the goal was not to prove the practicality of using ChatGPT for transportation, but to see if these models possessed enough spatial reasoning to operate a vehicle at low speeds. The team noted that they had to iterate on their prompts for hours to bypass the models' refusals to drive, eventually labeling the environment a "sandbox" to ensure the AI would comply.

