PulseAugur
中
实时 07:52:13
English(EN) Code as Representation: A Compilable Parsing Paradigm for Academic Documents

新的CADP范式旨在解锁学术论文中的知识

研究人员引入了一种名为可编译学术文档解析(CADP)的新范式,以更好地表示嵌入在学术论文中的科学知识。当前的方法在保留表格、公式和伪代码等元素的结构和逻辑方面存在困难,而这些元素对于多模态大语言模型(MLLMs)至关重要。CADP使用LaTeX和可执行Python重建这些文档,从而可以直接将重建的元素与源进行验证。一个新的基准CADP-Bench已被开发出来以评估此过程,结果显示即使是先进的MLLMs在生成高保真、可执行的重建方面仍有很大的改进空间。 AI

影响 这项研究可能有助于AI模型更有效地从科学文献中提取和利用知识。

排序理由 该条目描述了一篇介绍学术文档新解析范式和基准的新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的CADP范式旨在解锁学术论文中的知识

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目描述了一篇介绍学术文档新解析范式和基准的新研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
51 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Rihui Jin, Jun Wang, chengyuan zhu, Liang Mingyu, Yue Gao, Li Yunxuan, Kuicai Dong, Guilin Qi, Lin Ren, Yongrui Chen, Xinbang Dai, Jiaqi Li, Tongtong Wu, Gholamreza Haffari ·

    代码即表征:面向学术文档的可编译解析范式

    arXiv:2608.17550v1 Announce Type: cross Abstract: Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For Multimodal Large Language Models (MLLMs), the core …