NVIDIA has released Qwen3.8-2.4T-A95B, a Mixture-of-Experts (MoE) model with 2.4 trillion parameters and 95 billion active parameters. This model utilizes NVFP4 quantization with 4-bit precision, optimized for vLLM and SGLang on GB200/B300 hardware. It supports a context length of 1 million tokens. AI
IMPACT This release offers a highly efficient MoE model with a large context window, potentially improving performance and reducing resource needs for complex AI tasks.
RANK_REASON Model release from a major AI lab (NVIDIA). [lever_c_demoted from frontier_release: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →