Researchers have developed AutoData, an agentic system designed to automate the selection of pre-training data for large language models. This agent treats data selection as a heuristic engineering problem, searching directly over executable algorithms that utilize document features like lexical statistics and perplexity. AutoData iteratively refines these algorithms based on validation feedback from a proxy model, discovering a data selection recipe that surpasses human-designed curation pipelines within a single night of searching. The discovered method also demonstrates effectiveness at larger scales, improving downstream metrics. AI
IMPACT Automates a critical, labor-intensive part of LLM development, potentially accelerating research and deployment.
RANK_REASON The item is a research paper detailing a new method for data selection in LLM pre-training. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Autodata
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →