Researchers have introduced AMIGO, a new benchmark designed to evaluate agentic vision-language models (VLMs) in long-horizon, multi-image scenarios. Unlike previous evaluations that focused on single-image interactions, AMIGO challenges models to identify a hidden target image from a gallery by asking a series of attribute-focused Yes/No questions. The benchmark aims to assess a model's ability to select informative questions, track constraints across turns, and perform fine-grained discrimination as evidence accumulates. Initial evaluations on open-source VLMs revealed that success rates alone can be misleading, as models may achieve correct answers without robust evidence verification or by violating interaction protocols. AI
IMPACT This benchmark could drive development of more robust and interactive vision-language models capable of complex, multi-turn reasoning.
RANK_REASON The item describes a new benchmark for evaluating AI models, presented in an academic paper. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →