PulseAugur
EN
LIVE 08:00:34

New MissionBench benchmark reveals MLLMs struggle with aerial embodied tasks

Researchers have introduced MissionBench, a new benchmark designed to evaluate the zero-shot, mission-level capabilities of Multimodal Large Language Models (MLLMs) in aerial 3D environments. The benchmark includes 120 missions across five simulated environments and four task families, requiring agents to autonomously plan, navigate, and report outcomes using only egocentric observations and action history. Initial tests with 22 open- and closed-source MLLMs revealed that even the top-performing models succeeded on less than 35% of missions, significantly below human performance of 84.4%, indicating the substantial difficulty of multi-step embodied tasks. The study also observed that larger general-purpose models show improved zero-shot embodied capabilities, suggesting that scaling models can enhance performance in these complex aerial scenarios. AI

IMPACT Highlights significant gaps in current MLLM capabilities for complex, multi-step embodied tasks, guiding future research in aerial AI agents.

RANK_REASON The cluster describes a new benchmark and research paper evaluating MLLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New MissionBench benchmark reveals MLLMs struggle with aerial embodied tasks

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Suman Navaratnarajah, Taehyoung Kim, Jona Ruthardt, Ishaan Bhimwal, Ryousuke Yamada, Yannik Blei, Wolfram Burgard, Yuki M Asano ·

    Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

    arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction…