Researchers have developed AutoData, an agentic system designed to automate the selection of pre-training data for large language models. This agent treats data selection as a heuristic engineering problem, searching directly over executable algorithms that utilize document features like lexical statistics and perplexity. AutoData iteratively refines these algorithms based on validation feedback from a proxy model, discovering a data selection recipe that surpasses human-designed curation pipelines within a single night of searching. The discovered method also demonstrates effectiveness at larger scales, improving downstream metrics. AI
影响 Automates a critical, labor-intensive part of LLM development, potentially accelerating research and deployment.
排序理由 The item is a research paper detailing a new method for data selection in LLM pre-training. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Autodata
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
- scite Smart Citations
AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →