A new research paper has identified a significant amount of extremist speech within the Dolma dataset, a large open training corpus used for the OLMo series of models. Researchers developed a pipeline combining automated processing and expert verification to estimate the prevalence of such content. Their findings suggest that Dolma likely contains hundreds of thousands of documents with extremist material, including direct calls for violence, raising concerns about data curation and the pre-training of large language models. AI
IMPACT Highlights potential risks in LLM training data, prompting better data curation and safety measures.
RANK_REASON Research paper analyzing LLM training data for harmful content. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →