NVIDIA has released quantized versions of two large language models, DeepSeek-V4-Pro-0813 and Qwen3.8-2.4T-A95B, utilizing their NVFP4 quantization method. The DeepSeek model, with 1.65 trillion parameters, employs Hybrid Attention and runs on Nvidia's Blackwell B200 hardware via SGLang. The Qwen model, a 2.4 trillion parameter MoE model with 95 billion active parameters, supports a 1 million token context length and operates on GB200/B300 hardware using vLLM and SGLang. AI
IMPACT Enables deployment of massive LLMs on new hardware with improved efficiency and performance.
RANK_REASON NVIDIA is releasing quantized versions of frontier models (DeepSeek, Qwen) for their new hardware, which is a direct release from a major AI infrastructure provider. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →