A new benchmark called Chart2Code has been introduced to evaluate the chart understanding and code generation capabilities of large multimodal models (LMMs). This benchmark features a hierarchical structure with three levels of increasing difficulty, designed to reflect real-world user scenarios. Chart2Code comprises 2,023 tasks across 22 chart types and includes metrics for code correctness and visual fidelity. Initial testing on 25 state-of-the-art LMMs, including GPT-5, revealed significant challenges, with GPT-5 achieving an average score of only 0.57 on code evaluation and 0.22 on chart quality for editing tasks. AI
IMPACT This benchmark's difficulty highlights current limitations in multimodal reasoning, potentially guiding future LMM development towards better chart understanding and code generation.
RANK_REASON New academic paper introducing a novel benchmark for evaluating multimodal models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →