Researchers have developed SLAI T-Rex, a framework for optimizing the full-parameter post-training of trillion-parameter MoE models on Ascend SuperPOD infrastructure. This system achieved 34.22% Model FLOPs Utilization (MFU), a 2.93x improvement over baseline methods, while maintaining training stability. The framework was then used to create specialized Operations Research (OR) models using a dataset of 10K samples, with the resulting DeepSeek-V4-Flash model outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash on OR tasks. AI
IMPACT Demonstrates a pathway for efficient trillion-parameter model training and specialization on non-GPU hardware, potentially enabling more complex reasoning capabilities.
RANK_REASON The cluster describes a research paper detailing a new framework and optimization techniques for training large language models on specific hardware.
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →