A new benchmark called DroneCATS evaluates multimodal large language models (MLLMs) as agents for controlling drones. The study found that while smaller, open-source models can navigate effectively, they struggle with adhering to action protocols and correctly terminating tasks. Frontier models also exhibit issues, particularly in multi-drone coordination, where they may fail to differentiate between distinct views. The research highlights a gap between MLLMs' spatial perception capabilities and their ability to execute disciplined, goal-oriented actions, especially under onboard compute constraints. AI
IMPACT Highlights a critical gap in current MLLMs for embodied AI tasks, indicating a need for models that can reliably follow protocols and recognize task completion.
RANK_REASON Academic paper introducing a new benchmark and evaluation of existing models.
Read on Hugging Face Daily Papers →
- Claude Opus-5
- Cosmos3-Edge-2B
- DroneCATS
- DroneCATS-Agent
- Gemini 3.7 Flash
- GPT-5
- Qwen 3.5
- arXiv
- Hugging Face
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →