PulseAugur
EN
LIVE 17:05:30

inclusionAI releases Ling-3.0-flash model with official FP8 weights

inclusionAI has released its Ling-3.0-flash model, available in both BF16 and an official FP8 version, on Hugging Face. The model boasts 127.5 billion total parameters with 5.1 billion active parameters, featuring a fine-grained architecture with 512 experts. The FP8 version is notably smaller, around 128GB, making it more accessible for users with significant unified memory or multi-GPU setups. AI

IMPACT Makes a new large language model with an efficient FP8 version available for broader use.

RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=2 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

inclusionAI releases Ling-3.0-flash model with official FP8 weights

COVERAGE [2]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/derspenti ·

    inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vfdeek/inclusionailing30flash_weights_are_up_on_hugging/"> <img alt="inclusionAI/Ling-3.0-flash weights are up on Hugging Face — MIT, BF16 plus an official FP8" src="https://external-preview.redd.it/N3g5MjI3N…

  2. r/LocalLLaMA TIER_1 English(EN) · /u/-Cubie- ·

    inclusionAI/Ling-3.0-flash · Hugging Face

    <!-- SC_OFF --><div class="md"><p>The Ling-3.0-flash MoE is now open-weighted at 124B A5B params. I know the original announcements were before the Kimi K3, DeepSeek-V4-Flash and Qwen3.8 hype, but this model might still have a good niche for itself due to its sizing. </p> <p>Disc…