PulseAugur
中
实时 07:31:37
English(EN) From Charts to Code: A Hierarchical Benchmark for Multimodal Models

新的Chart2Code基准测试揭示GPT-5在多模态任务上存在困难

一项名为Chart2Code的新基准测试已被推出,用于评估大型多模态模型(LMM)的图表理解和代码生成能力。该基准测试具有层级结构,包含三个难度递增的级别,旨在反映真实用户场景。Chart2Code包含2,023个任务,涵盖22种图表类型,并包含代码正确性和视觉保真度的指标。对包括GPT-5在内的25个最先进LMM进行的初步测试显示出显著的挑战,GPT-5在代码评估上的平均得分仅为0.57,在编辑任务的图表质量上的得分仅为0.22。 AI

影响 该基准测试的难度凸显了当前多模态推理的局限性,可能指导未来LMM在图表理解和代码生成方面的发展。

排序理由 一项介绍用于评估多模态模型的新颖基准测试的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的Chart2Code基准测试揭示GPT-5在多模态任务上存在困难

本文如何被排名

Signal score
21 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
一项介绍用于评估多模态模型的新颖基准测试的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Jiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng, Min Li, Alex Jinpeng Wang ·

    从图表到代码:多模态模型的层级基准测试

    arXiv:2510.17932v5 Announce Type: replace-cross Abstract: We introduce Chart2Code, a new benchmark for evaluating the chart understanding and code generation capabilities of large multimodal models (LMMs). Chart2Code is explicitly designed from a user-driven perspective, capturin…