PulseAugur
实时 00:54:41
Deutsch(DE) NVIDIA stellt DeepSeek-V4-Pro-0813 als NVFP4-Quantisierung bereit. Das 1,65T-Parameter-MoE-Modell nutzt Hybrid-Attention und läuft auf Blackwell B200 via SGLang

NVIDIA 为 Blackwell 硬件发布量化版 DeepSeek 和 Qwen LLM

NVIDIA 已发布两种大型语言模型的量化版本,DeepSeek-V4-Pro-0813Qwen3.8-2.4T-A95B,采用了其 NVFP4 量化方法。DeepSeek 模型拥有 1.65 万亿参数,采用混合注意力机制,并通过 SGLang 在 Nvidia 的 Blackwell B200 硬件上运行。Qwen 模型是一个拥有 2.4 万亿参数、950 亿激活参数的 MoE 模型,支持 100 万 token 的上下文长度,并在 GB200/B300 硬件上使用 vLLM 和 SGLang 运行。 AI

影响 能够在大规模 LLM 上部署新硬件,提高效率和性能。

排序理由 NVIDIA 正在为其新硬件发布前沿模型(DeepSeek、Qwen)的量化版本,这是来自主要 AI 基础设施提供商的直接发布。[lever_c_demoted from frontier_release: ic=2 ai=1.0]

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

NVIDIA 为 Blackwell 硬件发布量化版 DeepSeek 和 Qwen LLM

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
NVIDIA 正在为其新硬件发布前沿模型(DeepSeek、Qwen)的量化版本,这是来自主要 AI 基础设施提供商的直接发布。[lever_c_demoted from frontier_release: ic=2 ai=1.0]
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
13 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [3]

  1. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA 提供 DeepSeek-V4-Flash-0731 作为 NVFP4 量化。该 304B MoE 模型使用 DSpark 进行推测解码,并通过 SGLang 在 B200 上运行。

    NVIDIA stellt DeepSeek-V4-Flash-0731 als NVFP4-Quantisierung bereit. Das 304B-MoE-Modell nutzt DSpark für spekulative Dekodierung und läuft auf B200 via SGLang/vLLM. MIT-Lizenz, 1M Kontext, Fokus auf Agenten und Tool-Use. https:// huggingface.co/nvidia/DeepSeek -V4-Flash-0731-NVF…

  2. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA 提供 DeepSeek-V4-Pro-0813 作为 NVFP4 量化。该 1.65T 参数 MoE 模型使用混合注意力,并通过 SGLang 在 Blackwell B200 上运行

    NVIDIA stellt DeepSeek-V4-Pro-0813 als NVFP4-Quantisierung bereit. Das 1,65T-Parameter-MoE-Modell nutzt Hybrid-Attention und läuft auf Blackwell B200 via SGLang. Die DSpark-Heads für spekulative Dekodierung bleiben unquantisiert. https:// huggingface.co/nvidia/DeepSeek -V4-Pro-08…

  3. Mastodon — mastodon.social TIER_1 Deutsch(DE) · aisyndicate ·

    NVIDIA 提供 Qwen3.8-2.4T-A95B 作为 NVFP4 量化。该 MoE 模型(2.4T 参数,95B 活跃)在 GB200/B300 上使用 4 位精度进行 vLLM/SGLang。

    NVIDIA stellt Qwen3.8-2.4T-A95B als NVFP4-Quantisierung bereit. Das MoE-Modell (2,4T Parameter, 95B aktiv) nutzt 4-Bit-Präzision für vLLM/SGLang auf GB200/B300. Kontextlänge: 1M Token. https:// huggingface.co/nvidia/Qwen3.8- 2.4T-A95B-NVFP4 # KI # AI # LLM # AISyndicate