Researchers have developed DEEPCHART, a new benchmark designed to evaluate the accuracy of Large Language Models (LLMs) in generating data-science charts. The benchmark, comprising 1,482 instances from real-world documents, assesses LLMs on their ability to extract relevant data, perform quantitative reasoning, and render charts faithfully. Experiments reveal that current state-of-the-art models often produce visually convincing charts that contain subtle data-level hallucinations, particularly in complex, multimodal contexts. The findings indicate that simply increasing context window sizes is not enough; reliable evidence extraction and reasoning capabilities are crucial for accurate chart generation. AI
IMPACT Highlights critical limitations in LLMs' ability to faithfully represent data visually, suggesting a need for improved data extraction and reasoning before chart rendering.
RANK_REASON The cluster describes a new academic paper introducing a benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →