PulseAugur
EN
LIVE 15:10:39

SLAI T-Rex framework optimizes DeepSeek-V4 models on Ascend SuperPOD

Researchers have developed SLAI T-Rex, a framework for optimizing the full-parameter post-training of trillion-parameter MoE models on Ascend SuperPOD infrastructure. This system achieved 34.22% Model FLOPs Utilization (MFU), a 2.93x improvement over baseline methods, while maintaining training stability. The framework was then used to create specialized Operations Research (OR) models using a dataset of 10K samples, with the resulting DeepSeek-V4-Flash model outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash on OR tasks. AI

IMPACT Demonstrates a pathway for efficient trillion-parameter model training and specialization on non-GPU hardware, potentially enabling more complex reasoning capabilities.

RANK_REASON The cluster describes a research paper detailing a new framework and optimization techniques for training large language models on specific hardware.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

SLAI T-Rex framework optimizes DeepSeek-V4 models on Ascend SuperPOD

COVERAGE [3]

  1. arXiv cs.AI TIER_1 English(EN) · Dongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen, Linyuan Qiu, Jian Meng, Zhengxuan Lu, Yiting Wang, Yucheng Xie, Tao Guo, Tianxiang Fang, Jing Li, Sihang Chen, Shihao Hong, Chang Liu, Weihua Dai, Zirong Zeng, Ziwei Zhu, Zhuohan Wang, Zhengjun Yue, Igor … ·

    SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

    arXiv:2607.20145v1 Announce Type: cross Abstract: Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and …

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

    Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-sca…

  3. r/LocalLLaMA TIER_1 English(EN) · /u/pmttyji ·

    [Paper] SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1v47kqc/paper_slai_trex_fullparameter_posttraining_of_the/"> <img alt="[Paper] SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD" src="https://preview.redd.it/3bpuadlysxeh1.…