A new research paper introduces a framework for evaluating interactive visual grounding in large vision-language models (LVLMs). The study highlights that current LVLMs struggle with tasks requiring dialogue to refine or acquire target information, performing significantly below human baselines. The research also found that LVLMs are poorly calibrated, with reported confidence often exceeding actual accuracy, indicating a need for further development in visual matching, information seeking, and synthesis capabilities. AI
影响 Highlights a significant gap in LVLM capabilities for interactive tasks, suggesting future research directions for more dynamic and context-aware AI systems.
排序理由 Academic paper introducing a new benchmark and evaluation framework for LVLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →