New Vision-Language-Action (VLA) models are emerging that aim to unify perception, reasoning, and control in autonomous driving, moving away from traditional modular stacks. Three prominent models—AutoVLA from UCLA, NVIDIA's Alpamayo family, and Qwen-Drive-1.0—represent different approaches to this unification. AutoVLA is designed for edge deployment with a small parameter count and token-based action output, while Alpamayo serves as a larger, cloud-scale teacher model intended for distillation into smaller runtime models. AI
IMPACT These unified VLA models could streamline autonomous driving systems, potentially leading to more efficient and capable self-driving vehicles.
RANK_REASON The article details new models and their architectures for autonomous driving, falling under research and development in AI. [lever_c_demoted from research: ic=1 ai=1.0]
- Alpamayo
- Autovla
- Autoware
- CARLA
- Cosmos 3 Super Reasoner
- Cosmos-Reason VLM
- nuPlan
- nuScenes
- NVIDIA
- Qwen2.5-VL-3B
- Qwen3-VL-32B
- Qwen-Drive-1.0
- UCLA
- Waymo E2E Challenge
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →