PulseAugur
实时 06:35:50
English(EN) Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training

新流程将科学论文转化为 LLM 的训练数据

研究人员开发了一种新流程,可将科学论文转化为多轮生成轨迹,用于语言模型的持续预训练。该方法重建了论文的写作过程,包括请求、计划和章节级别的讨论,同时保留原文 verbatim。由此产生的语料库来自 arXiv 论文,其大小约为源文本的两倍。这种方法还可以创建指令数据集和一个名为 PAW-Bench 的新学术写作基准。实验表明,在此语料库上进行持续预训练,然后进行监督微调,可以显著提高写作能力,而不会损害一般推理或长文档理解能力。 AI

影响 通过利用结构化的科学论文来增强 LLM 训练数据,有可能改善学术写作和长文档理解。

排序理由 该集群描述了一种处理科学论文以创建语言模型训练数据的新方法,该方法在 arXiv 预印本中有所介绍。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新流程将科学论文转化为 LLM 的训练数据

本文如何被排名

Signal score
29 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一种处理科学论文以创建语言模型训练数据的新方法,该方法在 arXiv 预印本中有所介绍。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Qiankai Xu, Qiguang Chen, Zixin Su, Wenhao Huang, Yue Gao, Jiaheng Liu, Ge Zhang ·

    将科学论文展开为多轮生成轨迹以进行持续预训练

    arXiv:2608.25826v1 Announce Type: new Abstract: A recent line of synthetic-data work reconstructs the thinking behind existing text rather than rewriting the text itself, but it operates on short web passages, recovers only local thoughts, and leaves the structure of whole docume…