A new benchmark called ChartAnno has been developed to evaluate how well multimodal large language models (MLLMs) can generate annotations for existing charts. The benchmark includes 1,200 real-world charts with associated code and annotation instructions, tested across different levels of specificity. Current proprietary models generally perform better, though large open-source models are closing the gap. The study found that more specific instructions enhance annotation quality, but inferring abstract intent remains a significant challenge for MLLMs, with chart images providing only marginal benefits. AI
IMPACT This benchmark highlights challenges in MLLMs' ability to generate chart annotations, indicating areas for future research and development in multimodal AI capabilities.
RANK_REASON The item describes a new academic benchmark for evaluating AI models. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →