PulseAugur
实时 16:06:59
English(EN) Characterizing Narrative Content in Web-scale LLM Pretraining Data

新框架分析 LLM 预训练数据中的叙事结构 · 跟踪 4 个来源

研究人员开发了一个新框架和模型 NarraBERT,用于分析大型语言模型 (LLM) 预训练数据中的叙事结构。该研究将此框架应用于 3 万亿 token 的 Dolma 语料库,创建了一个名为 NarraDolma 的新数据集。研究结果表明,叙事质量在各种数据来源和主题中分布不均,这表明当前的语料库构建实践并未考虑到这些细微差别。发布的框架、数据集和模型旨在为理解叙事数据组成及其对 LLM 推理的影响奠定基础。 AI

影响 提供了工具和见解,用于理解训练数据中的叙事质量可能如何影响 LLM 的行为和推理。

排序理由 该集群包含两篇详细介绍 LLM 预训练数据和叙事分析研究的学术论文,其中包括发布的新模型和数据集。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新框架分析 LLM 预训练数据中的叙事结构 · 跟踪 4 个来源

报道来源 [3]

  1. arXiv cs.CL TIER_1 English(EN) · Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak ·

    网络规模LLM预训练数据中叙事内容的特征分析

    arXiv:2606.19468v1 Announce Type: new Abstract: The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a …

  2. arXiv cs.AI TIER_1 English(EN) · David Y. Liu, Aditya Joshi, Paul Dawson ·

    面向故事自动生成与理解的叙事理论驱动大语言模型方法:一项调查

    arXiv:2602.15851v2 Announce Type: replace-cross Abstract: Applications of narrative theories using large language models (LLMs) deliver promising methods in automatic story generation and understanding tasks. Our survey examines how natural language processing (NLP) research uses…

  3. arXiv cs.CL TIER_1 English(EN) · Maria Antoniak ·

    表征网络规模LLM预训练数据中的叙事内容

    The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token open pretraining corpus. Drawin…