PulseAugur
中
实时 15:41:32
English(EN) Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs

Nereus运行时实时调整大型语言模型训练后并行策略

研究人员开发了Nereus,一个新颖的运行时系统,旨在实时调整大型语言模型(LLM)训练后任务的并行策略。当资源可用性或内存压力等因素在运行时发生变化时,传统方法难以应对,导致效率低下。Nereus通过动态选择和切换到新的、优化的执行计划来解决这个问题,将分布式模型状态表示为“弹性模型单元”,并使用过渡图来管理GPU传输。这种自适应方法已显示出显著的改进,与OpenRLHF等现有系统相比,平均步长延迟降低了27.7%,端到端吞吐量提高了7.27倍。 AI

影响 这种自适应运行时可以显著提高训练大型语言模型的效率并降低成本。

排序理由 该集群描述了一篇详细介绍LLM训练后新系统的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 7 个来源。 我们如何撰写摘要 →

Nereus运行时实时调整大型语言模型训练后并行策略

本文如何被排名

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇详细介绍LLM训练后新系统的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
infra, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
2 days old
Coverage has settled into its steady-state source set.
Coverage growth since scoring
+6 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [7]

  1. arXiv cs.AI TIER_1 English(EN) · Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen, Ziqi Liu, Zhiqiang Zhang, Xiaolin Huang, Jun Zhou ·

    轨迹汤:通过多样化轨迹推动 LLM 中期训练的计算扩展前沿

    arXiv:2609.37169v1 Announce Type: cross Abstract: Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even …

  2. arXiv cs.AI TIER_1 English(EN) · Chenliang Li, Neiwen Ling, Zijun Wei, Alfredo Garcia ·

    停滞在哪里累积?用于 LLM 训练后异步 RL 的池感知有效停滞控制

    arXiv:2609.36830v1 Announce Type: new Abstract: Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are gen…

  3. arXiv cs.AI TIER_1 English(EN) · Pengyu He, Yan Zhang, Ruien Li, Guangwen Yang ·

    AutoLoCo:通过自适应同步实现通信高效的分布式大模型训练

    arXiv:2609.36662v1 Announce Type: cross Abstract: The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fractio…

  4. arXiv cs.LG TIER_1 English(EN) · Haozhan Tang, Hao Kang, Han Cai, Song Han, Chenyan Xiong ·

    Delta-Matching:弥合大语言模型原生8位训练的最后一道鸿沟

    arXiv:2609.37852v1 Announce Type: new Abstract: Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta…

  5. arXiv cs.CL TIER_1 English(EN) · Tianhao Qian, Ziming Hong, Chongyang Gao, Kezhen Chen, Lixu Wang ·

    存储并非策略:LLM 遗忘的状态条件支持控制

    arXiv:2609.37858v1 Announce Type: cross Abstract: Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the …

  6. arXiv cs.LG TIER_1 English(EN) · Shufan Shen, Zhongni Hou, Junshu Sun, Yufei Zhang, Wei Lin, Guojun Yin, Qingming Huang, Shuhui Wang ·

    LLM训练后泛化能力源于策略内参数更新方向

    arXiv:2609.36659v1 Announce Type: new Abstract: The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, over…

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    Nereus:LLM 训练后自适应并行

    Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage…