DriveMA: Rethinking Language Interfaces in Driving VLAs with One-Step Meta-Actions
Researchers have introduced DriveMA, a new approach for driving vision-language-action models that replaces complex natural language reasoning with simpler, one-step meta-actions. This method addresses bottlenecks in annotation, model complexity, and inference latency associated with traditional reasoning-centric interfaces. DriveMA achieves new state-of-the-art results on the Waymo End-to-End Driving Challenge, demonstrating the effectiveness of its action-centric supervised training and reinforcement learning framework. AI
IMPACT Simplifies driving AI interfaces, potentially improving efficiency and scalability for autonomous vehicle development.