Skip to content
Breaking:

Engineers Prompt General-Purpose AI Models to Steer a Car in Real-World Test

An independent experiment using OpenAI's GPT-6 Astra to navigate a Toyota Corolla highlights emerging physical reasoning in multimodal models.

By The Company Wire3 min read
Share
OpenAI — Engineers Prompt General-Purpose AI Models to Steer a Car in Real-World Test
OpenAI — Engineers Prompt General-Purpose AI Models to Steer a Car in Real-World Test. Photo: Wired.

Three artificial intelligence engineers recently tested the real-world physical reasoning of general-purpose language models by using OpenAI's GPT-6 Astra to steer a 2024 Toyota Corolla through a Bay Area fast-food drive-thru lane. Aditya Ramabadran, Simon Mahns, and Tobias Gessler—engineers at startup Axiom conducting an independent weekend project—connected a laptop running a chat interface to a server linked to windshield cameras and the vehicle's power steering, as reported by Wired (https://www.wired.com/story/ai-is-driving-cars-now-oh-boy/).

With a safety driver keeping a foot over the brake pedal, GPT-6 Astra slowly guided the car up to the In-N-Out Burger pickup window. Unlike dedicated autonomous driving stacks engineered specifically for vehicular navigation by companies like Tesla and Waymo, the demonstration relied on a foundation model originally developed for text, code, and multimodal tasks, with no prior vehicular training by the researchers.

The engineers began testing vehicular navigation after noticing that recent models had grown proficient at generating complex 3D simulations. When they initially tested SpaceXAI's Grok alongside models from OpenAI and Anthropic, the systems declined direct control commands citing safety limitations. After prompt adjustments bypassed those initial refusals, the models began issuing steering inputs. The team noted that the systems appeared to adapt via in-context learning to adjust control mechanics in response to errors.

To measure spatial navigation systematically, the trio created an evaluation benchmark named DrivingBench, which tests steering across a simplified parking lot course. GPT-6 Astra was the only tested model to complete the full course, though it moved at a very low speed. Anthropic's Claude Fable 5.1 completed 45 percent of the course, while SpaceXAI's Grok completed 11 percent.

The project coincides with broader industry efforts to advance physical scene comprehension beyond standard image recognition. Andrew Dai, chief executive of Elorian AI and a former Google DeepMind researcher, noted that physical reasoning is a necessary prerequisite for practical applications such as home robotics. Elorian AI and Scale AI recently launched "Humanity's Sixth Sense," a benchmark developed to assess how intuitively machine learning models perceive and interpret physical environments.

Sources

  1. Wired

Company: OpenAI

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.