PulseAugur
EN
LIVE 19:43:07
Deutsch(DE) NVIDIA stellt Qwen3.8-2.4T-A95B als NVFP4-Quantisierung bereit. Das MoE-Modell (2,4T Parameter, 95B aktiv) nutzt 4-Bit-Präzision für vLLM/SGLang auf GB200/B300.

NVIDIA releases Qwen3.8-2.4T-A95B MoE model with 4-bit quantization

NVIDIA has released Qwen3.8-2.4T-A95B, a Mixture-of-Experts (MoE) model with 2.4 trillion parameters and 95 billion active parameters. This model utilizes NVFP4 quantization with 4-bit precision, optimized for vLLM and SGLang on GB200/B300 hardware. It supports a context length of 1 million tokens. AI

IMPACT This release offers a highly efficient MoE model with a large context window, potentially improving performance and reducing resource needs for complex AI tasks.

RANK_REASON Model release from a major AI lab (NVIDIA). [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

NVIDIA releases Qwen3.8-2.4T-A95B MoE model with 4-bit quantization

How we ranked this

Signal score
14 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
Model release from a major AI lab (NVIDIA). [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA provides Qwen3.8-2.4T-A95B as NVFP4 quantization. The MoE model (2.4T parameters, 95B active) uses 4-bit precision for vLLM/SGLang on GB200/B300.

    NVIDIA stellt Qwen3.8-2.4T-A95B als NVFP4-Quantisierung bereit. Das MoE-Modell (2,4T Parameter, 95B aktiv) nutzt 4-Bit-Präzision für vLLM/SGLang auf GB200/B300. Kontextlänge: 1M Token. https:// huggingface.co/nvidia/Qwen3.8- 2.4T-A95B-NVFP4 # KI # AI # LLM # AISyndicate