English

Model ReleasesAxiom MathGPT-6 Astra

Frontier LLMs struggle to drive real cars, with GPT-6 Astra being the only to complete a course

A technical project named DrivingBench has conducted experiments to test whether frontier general-purpose large language models (LLMs) can operate a real vehicle in the physical world. The experiment, conducted by developers from Axiom Math, utilized a 2022 Toyota Corolla equipped with a "comma four" device and openpilot software to bridge the AI's commands to the vehicle's controls.

The setup allowed models to control steering, speed, and duration through an MCP (Model Context Protocol) server. The models were tasked with navigating a cone-lined course in a parking lot, including a final parking maneuver.

The results showed varying levels of success. GPT-6 Astra was the only model to successfully complete the course, finishing a 135m course in 5 minutes and 22 seconds on its second attempt. Claude Fable 5.1 managed to progress through approximately 45% of the course, while Grok 4.6 and GPT-5.6 Sol failed to move past the initial turn in all attempts.

The researchers reported significant technical challenges during the testing. Interestingly, several models, including GPT-6 Astra, initially refused to operate the car when they recognized the input was a real vehicle rather than a simulation, citing safety concerns. The team was only able to achieve stable operation by renaming the MCP server to "DrivingBench Sandbox."

Beyond navigation accuracy, the experiment highlighted "response latency" as a critical factor. Because the vehicle continues to move based on the last received command while the model is processing the next step, delays in model reasoning increase the risk of the vehicle veering off course. The team plans to develop DrivingBench v2 to evaluate more difficult courses and increase the number of trials per model.

Sources

  1. ChatGPTなどの汎用AIに本物の車を運転させたら「実車では?」と気付いて運転を拒否、最終的に走らせた方法とは (GIGAZINE, 2026-09-30)
  2. DrivingBench