PulseAugur
实时 07:25:07
English(EN) Pair-Level Essay-Scale Republication and Reuse from Fragmented Historical Text Reuse: A Workflow Study on Eighteenth-Century Books and Newspapers

新工作流程借助大语言模型识别历史文本再利用

研究人员开发了一种新的工作流程,用于识别文章级别的重印和碎片化历史文本的再利用,重点关注十八世纪哲学家 David Hume 的作品。该研究将分阶段的基于规则的工作流程与直接的大语言模型设置和自动化规则改编进行了比较。所提出的工作流程在标记数据上取得了较高的 F1 分数,并展示了强大的精确率-召回率权衡,有效地将证据整合为合理的传播关系。该方法为创建用于历史分析的紧凑候选空间提供了一种实用的方法,即使在地面真实信息不完整的情况下也是如此。 AI

影响 这项研究展示了大语言模型在历史文本分析方面的新应用,有望提高数字人文研究的准确性和效率。

排序理由 这是一篇详细介绍文本再利用分析新方法的学术论文。[lever_c_demoted from research: ic=1 ai=0.7]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新工作流程借助大语言模型识别历史文本再利用

本文如何被排名

Signal score
16 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是一篇详细介绍文本再利用分析新方法的学术论文。[lever_c_demoted from research: ic=1 ai=0.7]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Ke Shu, Kira Hinderks, Eetu M\"akel\"a, Mikko Tolonen ·

    从碎片化历史文本重用中进行成对文章级重版和再利用:一项关于十八世纪书籍和报纸的工作流程研究

    arXiv:2608.27343v1 Announce Type: new Abstract: This paper addresses the recovery of essay-scale republication and reuse from fragmented text-reuse evidence, a setting whose central challenge is pair-level evidence consolidation and not fragment retrieval alone. The study focuses…