Researchers have introduced MissionBench, a new benchmark designed to evaluate the zero-shot, mission-level capabilities of Multimodal Large Language Models (MLLMs) in aerial 3D environments. The benchmark includes 120 missions across five simulated environments and four task families, requiring agents to autonomously plan, navigate, and report outcomes using only egocentric observations and action history. Initial tests with 22 open- and closed-source MLLMs revealed that even the top-performing models succeeded on less than 35% of missions, significantly below human performance of 84.4%, indicating the substantial difficulty of multi-step embodied tasks. The study also observed that larger general-purpose models show improved zero-shot embodied capabilities, suggesting that scaling models can enhance performance in these complex aerial scenarios. AI
IMPACT Highlights significant gaps in current MLLM capabilities for complex, multi-step embodied tasks, guiding future research in aerial AI agents.
RANK_REASON The cluster describes a new benchmark and research paper evaluating MLLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →