PulseAugur
中
实时 17:30:51
English(EN) From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs

新基准揭示多模态大语言模型在视觉到代码复现方面存在困难

一个名为 FigCodeBench 的新基准已被开发出来,用于评估多模态大语言模型(MLLMs)复现复杂视觉图形和生成相应代码的能力。该框架通过整合视觉理解和代码生成,超越了孤立的评估,解决了当前基准的空白。在包括 Gemini-3.1 Pro 和 GPT-5.4 在内的 24 个专有和开源 MLLMs 上进行的实验,揭示了它们在各种编程语言和难度级别上的显著性能下降,为理解它们的局限性提供了见解。 AI

影响 凸显了当前多模态大语言模型在复杂视觉到代码任务方面的局限性,可能指导未来的模型开发。

排序理由 介绍多模态大语言模型新基准和评估框架的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准揭示多模态大语言模型在视觉到代码复现方面存在困难

本文如何被排名

Signal score
4 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
介绍多模态大语言模型新基准和评估框架的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Zijian Chen, Zhengyu Chen, Bohan Liang, Lirong Deng, Yushuo Zheng, Yanwei Jiang, Qi Jia, Kaiwei Zhang, Wenjun Zhang, Guangtao Zhai ·

    从像素到编码:评估多模态大语言模型图像复现能力

    arXiv:2610.10066v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedica…