PulseAugur
EN
LIVE 23:37:54

Huawei open-sources openPangu-2.0-Pro MoE model with 505B parameters

Huawei has open-sourced its new Mixture-of-Experts (MoE) model, openPangu-2.0-Pro. This model boasts a total of 505 billion parameters with 18 billion activated, and supports a context length of 512,000 tokens. It was trained on 34 trillion tokens and incorporates advanced training techniques including unified SFT for slow and fast thinking, multiple specialist RL training, and on-policy distillation. AI

IMPACT This open-source release provides researchers and developers with a powerful new MoE model, potentially accelerating advancements in efficient large language model development and deployment.

RANK_REASON Open-source release of a large language model from a major tech company. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Huawei open-sources openPangu-2.0-Pro MoE model with 505B parameters

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/langsfang ·

    Huawei opensouced openPangu-2.0-Pro, 505B-A18B

    <!-- SC_OFF --><div class="md"><p>openPangu-2.0-Pro is an MoE model trained on Ascend. The model has 505B total parameters and 18B activated parameters. Its context length is 512k. The total pretraining data contains 34T tokens. During Post-training, openPangu-2.0-Pro is trained …