Researchers have developed a new framework and model, NarraBERT, to analyze narrative structures within large language model (LLM) pretraining data. The study applied this framework to the 3-trillion-token Dolma corpus, creating a new dataset called NarraDolma. Findings indicate that narrative qualities are unevenly distributed across various data sources and topics, suggesting current data curation practices do not account for these nuances. The released framework, dataset, and model aim to provide a foundation for understanding narrative data composition and its impact on LLM reasoning. AI
IMPACT Provides tools and insights for understanding how narrative qualities in training data might influence LLM behavior and reasoning.
RANK_REASON The cluster contains two academic papers detailing research into LLM pretraining data and narrative analysis, including the release of a new model and dataset.
Read on Hugging Face Daily Papers →
- alphaXiv
- arXiv
- DagsHub
- David T. Liu
- Gotit.pub
- Hugging Face
- large language models
- narratology
- natural language processing
- ScienceCast
- Dolma
- NarraBERT
- NarraDolma
- Roberta
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →