DCLM
PulseAugur coverage of DCLM — every cluster mentioning DCLM across labs, papers, and developer communities, ranked by signal.
-
New CuraWeb corpus boosts LLM performance with optimized data curation
Researchers have developed CuraWeb, a new 2 trillion token English corpus designed to improve the pretraining data for large language models. Unlike previous methods that focused on singular optimization objectives, Cur…
-
AI Models Struggle with Copying and Visualized Text, New Research Shows
Two new research papers highlight limitations in current AI models. One paper, "Frontier Language Models Struggle to Copy," reveals that even advanced large language models fail at simple string copying tasks due to the…
-
Scaling LLMs improves social simulations, but with limitations
A new research paper explores the impact of scaling Large Language Models (LLMs) on their ability to perform social simulations. The study found that increasing the compute scale of LLMs, specifically using the Qwen3 ar…
-
Spokes framework boosts AI pretraining data diversity by 489%
Researchers have developed a new probabilistic diversification framework called Spokes, which optimizes for diversity in pretraining data selection. This method utilizes the G-Vendi score and exponentiated gradient desc…
-
Researchers track attention circuit formation in 1B-class language models
A new research paper investigates the emergence of attention circuits in language models, specifically tracking how different types of attention heads form across various model architectures and training datasets. The s…