English(EN)Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
Nereus运行时实时调整大型语言模型训练后并行策略
作者PulseAugur 编辑部·[7 个来源]·
研究人员开发了Nereus,一个新颖的运行时系统,旨在实时调整大型语言模型(LLM)训练后任务的并行策略。当资源可用性或内存压力等因素在运行时发生变化时,传统方法难以应对,导致效率低下。Nereus通过动态选择和切换到新的、优化的执行计划来解决这个问题,将分布式模型状态表示为“弹性模型单元”,并使用过渡图来管理GPU传输。这种自适应方法已显示出显著的改进,与OpenRLHF等现有系统相比,平均步长延迟降低了27.7%,端到端吞吐量提高了7.27倍。
AI
影响
这种自适应运行时可以显著提高训练大型语言模型的效率并降低成本。
排序理由
该集群描述了一篇详细介绍LLM训练后新系统的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
arXiv:2609.37169v1 Announce Type: cross Abstract: Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even …
arXiv:2609.36830v1 Announce Type: new Abstract: Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are gen…
arXiv cs.AI
TIER_1English(EN)·Pengyu He, Yan Zhang, Ruien Li, Guangwen Yang·
arXiv:2609.36662v1 Announce Type: cross Abstract: The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fractio…
arXiv cs.LG
TIER_1English(EN)·Haozhan Tang, Hao Kang, Han Cai, Song Han, Chenyan Xiong·
arXiv:2609.37852v1 Announce Type: new Abstract: Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta…
arXiv:2609.37858v1 Announce Type: cross Abstract: Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the …
arXiv:2609.36659v1 Announce Type: new Abstract: The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, over…
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage…