PulseAugur
实时 09:34:46
English(EN) Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

新管道从历史报纸中提取数十亿词元

研究人员开发了机构报纸管道(Institutional Newspapers Pipeline),这是一个模块化系统,旨在从历史报纸扫描件中提取高质量、结构化数据。该管道与波士顿公共图书馆合作开发,通过分割、OCR和各种文本分析步骤处理扫描件,以分类类型、检测阅读顺序、识别命名实体并生成嵌入。该系统应用于1795年至1930年间发布的超过140万份公共领域报纸扫描件,产生了163亿词元,并形成了一个开放数据集。 AI

影响 能够对历史文本数据进行大规模分析,有可能从数字化的档案中发现新的见解。

排序理由 该项目是一篇研究论文,描述了一个新的数据提取管道。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新管道从历史报纸中提取数十亿词元

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Matteo Cargnelutti, Catherine Brobston, Eben English, Jake Sadow, Kacie Bailey, Greg Leppert, Amanda Watson, Jessica Chapel, Jonathan Zittrain ·

    机构报纸管道:从历史报纸中提取数十亿高质量词元

    arXiv:2608.18972v1 Announce Type: new Abstract: Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts make computational access to these materials both challenging and limited. We present the Institutional Newspapers P…