PulseAugur
中
实时 23:29:52
English(EN) AutoData: Agentic Search for Pre-training Data Selection

新的代理式系统可自动选择用于大型语言模型的预训练数据

研究人员开发了AutoData,一个代理式系统,旨在自动选择用于大型语言模型的预训练数据。该代理将数据选择视为一个启发式工程问题,直接搜索利用词汇统计和困惑度等文档特征的可执行算法。AutoData根据代理模型的验证反馈迭代地改进这些算法,在单晚搜索中发现了一种数据选择方法,该方法优于人类设计的策展流程。所发现的方法在大规模应用中也显示出有效性,并改进了下游指标。 AI

影响 自动化了大型语言模型开发中一个关键的、劳动密集型的部分,有可能加速研究和部署。

排序理由 该项目是一篇研究论文,详细介绍了大型语言模型预训练中数据选择的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的代理式系统可自动选择用于大型语言模型的预训练数据

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目是一篇研究论文,详细介绍了大型语言模型预训练中数据选择的新方法。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
20 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Yan Meng, Dhruv Srikanth, Bingchen Zhao, Zhengyao Jiang, Yuxiang Wu ·

    AutoData:用于预训练数据选择的代理搜索

    arXiv:2609.19754v1 Announce Type: new Abstract: LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-train…