PulseAugur
中
实时 04:12:45
English(EN) POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction

轻量级POTATR模型实现SOTA表格提取

研究人员开发了POTATR,一种用于从文档中提取表格的新型轻量级图像到图模型。这个拥有2900万参数的模型在PubTables-v2基准测试上显著优于现有方法,取得了0.964的GriTS_Con分数。与当前的大型语言模型相比,POTATR的速度更快、成本效益更高,其输出具有空间定位功能,便于验证和进一步集成。 AI

影响 为高效准确的表格提取设定了新标准,有望加速文档处理工作流程。

排序理由 关于新模型和基准测试结果的学术论文。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

轻量级POTATR模型实现SOTA表格提取

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
关于新模型和基准测试结果的学术论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
122 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.CV TIER_1 English(EN) · Brandon Smock, Libin Liang, Max Sokolov, Amrit Ramesh, Valerie Faucon-Morin, Tayyibah Khanam, Maury Courtland ·

    POTATR:一种用于页面级表格提取的轻量级图像到图模型

    arXiv:2606.09788v1 Announce Type: new Abstract: Large-scale document processing requires contextually aware table extraction (TE) that is both accurate and efficient. Yet current approaches require billions of parameters, hundreds of autoregressive steps, or costly API inference.…

  2. arXiv cs.CV TIER_1 English(EN) · Maury Courtland ·

    POTATR:一种用于页面级表格提取的轻量级图像到图模型

    Large-scale document processing requires contextually aware table extraction (TE) that is both accurate and efficient. Yet current approaches require billions of parameters, hundreds of autoregressive steps, or costly API inference. Motivated by this, we introduce the Page-Object…