PulseAugur
EN
LIVE 13:14:18

New framework analyzes narrative structure in LLM pretraining data · 4 sources tracked

Researchers have developed a new framework and model, NarraBERT, to analyze narrative structures within large language model (LLM) pretraining data. The study applied this framework to the 3-trillion-token Dolma corpus, creating a new dataset called NarraDolma. Findings indicate that narrative qualities are unevenly distributed across various data sources and topics, suggesting current data curation practices do not account for these nuances. The released framework, dataset, and model aim to provide a foundation for understanding narrative data composition and its impact on LLM reasoning. AI

IMPACT Provides tools and insights for understanding how narrative qualities in training data might influence LLM behavior and reasoning.

RANK_REASON The cluster contains two academic papers detailing research into LLM pretraining data and narrative analysis, including the release of a new model and dataset.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

New framework analyzes narrative structure in LLM pretraining data · 4 sources tracked

COVERAGE [3]

  1. arXiv cs.CL TIER_1 English(EN) · Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak ·

    Characterizing Narrative Content in Web-scale LLM Pretraining Data

    arXiv:2606.19468v1 Announce Type: new Abstract: The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a …

  2. arXiv cs.AI TIER_1 English(EN) · David Y. Liu, Aditya Joshi, Paul Dawson ·

    Narrative Theory-Driven LLM Methods for Automatic Story Generation and Understanding: A Survey

    arXiv:2602.15851v2 Announce Type: replace-cross Abstract: Applications of narrative theories using large language models (LLMs) deliver promising methods in automatic story generation and understanding tasks. Our survey examines how natural language processing (NLP) research uses…

  3. arXiv cs.CL TIER_1 English(EN) · Maria Antoniak ·

    Characterizing Narrative Content in Web-scale LLM Pretraining Data

    The narrative composition of web-scale LLM pretraining corpora remains largely unexplored even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token open pretraining corpus. Drawin…