PulseAugur
EN
LIVE 00:18:30
Deutsch(DE) NVIDIA stellt DeepSeek-V4-Pro-0813 als NVFP4-Quantisierung bereit. Das 1,65T-Parameter-MoE-Modell nutzt Hybrid-Attention und läuft auf Blackwell B200 via SGLang

NVIDIA releases quantized DeepSeek and Qwen LLMs for Blackwell hardware

NVIDIA has released quantized versions of two large language models, DeepSeek-V4-Pro-0813 and Qwen3.8-2.4T-A95B, utilizing their NVFP4 quantization method. The DeepSeek model, with 1.65 trillion parameters, employs Hybrid Attention and runs on Nvidia's Blackwell B200 hardware via SGLang. The Qwen model, a 2.4 trillion parameter MoE model with 95 billion active parameters, supports a 1 million token context length and operates on GB200/B300 hardware using vLLM and SGLang. AI

IMPACT Enables deployment of massive LLMs on new hardware with improved efficiency and performance.

RANK_REASON NVIDIA is releasing quantized versions of frontier models (DeepSeek, Qwen) for their new hardware, which is a direct release from a major AI infrastructure provider. [lever_c_demoted from frontier_release: ic=2 ai=1.0]

Read on Mastodon — mastodon.social →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

NVIDIA releases quantized DeepSeek and Qwen LLMs for Blackwell hardware

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
NVIDIA is releasing quantized versions of frontier models (DeepSeek, Qwen) for their new hardware, which is a direct release from a major AI infrastructure provider. [lever_c_demoted from frontier_…
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
13 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [3]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA provides DeepSeek-V4-Flash-0731 as NVFP4 quantization. The 304B MoE model uses DSpark for speculative decoding and runs on B200 via SGLang.

    NVIDIA stellt DeepSeek-V4-Flash-0731 als NVFP4-Quantisierung bereit. Das 304B-MoE-Modell nutzt DSpark für spekulative Dekodierung und läuft auf B200 via SGLang/vLLM. MIT-Lizenz, 1M Kontext, Fokus auf Agenten und Tool-Use. https:// huggingface.co/nvidia/DeepSeek -V4-Flash-0731-NVF…

  2. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA provides DeepSeek-V4-Pro-0813 as NVFP4 quantization. The 1.65T parameter MoE model uses Hybrid-Attention and runs on Blackwell B200 via SGLang

    NVIDIA stellt DeepSeek-V4-Pro-0813 als NVFP4-Quantisierung bereit. Das 1,65T-Parameter-MoE-Modell nutzt Hybrid-Attention und läuft auf Blackwell B200 via SGLang. Die DSpark-Heads für spekulative Dekodierung bleiben unquantisiert. https:// huggingface.co/nvidia/DeepSeek -V4-Pro-08…

  3. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA provides Qwen3.8-2.4T-A95B as NVFP4 quantization. The MoE model (2.4T parameters, 95B active) uses 4-bit precision for vLLM/SGLang on GB200/B300.

    NVIDIA stellt Qwen3.8-2.4T-A95B als NVFP4-Quantisierung bereit. Das MoE-Modell (2,4T Parameter, 95B aktiv) nutzt 4-Bit-Präzision für vLLM/SGLang auf GB200/B300. Kontextlänge: 1M Token. https:// huggingface.co/nvidia/Qwen3.8- 2.4T-A95B-NVFP4 # KI # AI # LLM # AISyndicate