A new benchmark called OVEarth-Bench has been introduced to evaluate open-vocabulary Earth observation capabilities. This benchmark addresses limitations in existing evaluations by expanding category breadth and query diversity, supporting mask and box localization under a zero-shot protocol. Initial evaluations indicate that current methods have limited performance, with multimodal large language models (MLLMs) showing the strongest results, while Earth observation-specific methods tend to underperform general models. AI
IMPACT This benchmark aims to guide the development of more realistic and diverse evaluation methods for open-vocabulary Earth observation, potentially improving MLLM performance in this area.
RANK_REASON The cluster describes a new academic benchmark for evaluating AI models in a specific domain.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →