PulseAugur
中
实时 20:29:19
English(EN) vllm v0.31.0 introduces major serving upgrades, making FlashMLA mega attention with V4.1 NVFP4 compressed KV cache the SM100 default. Release includes DeepGEMM

vLLM v0.31.0 通过 FlashMLA、DeepGEMM 和 MXFP8 增强服务

vLLM 发布了 0.31.0 版本,为其服务能力带来了重大增强。此次更新将 FlashMLA 与 V4.1 NVFP4 压缩 KV 缓存集成,作为 SM100 的默认配置,并包含了 DeepGEMM。该版本还具有稀疏 MQA logits、Mega-Gate 融合以及使用 MXFP8 量化的融合小批量 WO-A。 AI

影响 vLLM 的最新版本为 AI 服务基础设施提供了性能改进。

排序理由 这是一个基础设施工具的软件发布,不是前沿模型发布或重大行业事件。

在 Mastodon — sigmoid.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

vLLM v0.31.0 通过 FlashMLA、DeepGEMM 和 MXFP8 增强服务

本文如何被排名

Signal score
3 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是一个基础设施工具的软件发布,不是前沿模型发布或重大行业事件。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    vllm v0.31.0 引入重大服务升级,使 FlashMLA 成为默认的 V4.1 NVFP4 压缩 KV 缓存 SM100 的 mega attention。发布包含 DeepGEMM

    vllm v0.31.0 introduces major serving upgrades, making FlashMLA mega attention with V4.1 NVFP4 compressed KV cache the SM100 default. Release includes DeepGEMM sparse MQA logits, Mega-Gate fusing, and fused small-batch WO-A with MXFP8 quant. # AI # MachineLearning # DevOps