PulseAugur
实时 06:11:41
English(EN) Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

新方法高效估计LLM超参数缩放定律

研究人员开发了一种名为幂律熵搜索(PLES)的新方法,以更有效地估计大型语言模型(LLM)的超参数缩放定律。该方法利用多保真度贝叶斯优化,侧重于减少缩放定律估计的整体不确定性,而不是优化单一目标。PLES 选择最大化每单位计算成本不确定性降低的候选配置,优先考虑信息量大的小规模实验。在合成数据、代理模型和实际 LLM 预训练运行上的评估表明,PLES 以传统网格搜索方法所需计算预算的十分之一不到的成本实现了准确的缩放定律。 AI

影响 该方法可以显著降低调整 LLM 的计算成本,从而可能加速研究和开发。

排序理由 学术论文,详细介绍了一种估计 LLM 超参数缩放定律的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新方法高效估计LLM超参数缩放定律

本文如何被排名

Signal score
34 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,详细介绍了一种估计 LLM 超参数缩放定律的新方法。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Eytan Bakshy, Jihao Andreas Lin ·

    通过幂律熵搜索高效估计最优超参数缩放定律

    arXiv:2609.01431v1 Announce Type: cross Abstract: Optimal hyperparameter scaling laws describe how the best hyperparameters for large language model (LLM) training change with model and data scale, enabling practitioners to predict optimal configurations at production scales with…