A new benchmark called DrivingBench has been developed to test the real-world driving capabilities of vision-language models. The benchmark uses a Toyota Corolla equipped with tools that allow models to see camera feeds and directly control the vehicle's steering and velocity. In initial tests, GPT-6 Astra was the only model among GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Grok 4.6 to successfully complete the course, demonstrating the current limitations of AI in performing complex, real-world tasks that require continuous action and monitoring. AI
IMPACT Tests the real-world applicability of vision-language models beyond digital benchmarks, highlighting current limitations in complex, long-horizon tasks.
RANK_REASON The item describes a new benchmark and research paper for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Claude Code
- Claude Fable 5.1
- Codex
- Cursor+
- DrivingBench
- GPT-5.6 Sol
- GPT-6 Astra
- Grok 4.6
- Toyota Corolla
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →