PulseAugur
中
实时 22:18:41
English(EN) v0.31.1rc0: [Metrics] Expose cached prompt tokens by cache tier (#56318)

vLLM 发布 v0.31.1rc0,新增缓存指标

vLLM 发布了 0.31.1rc0 版本,引入了一项新指标,用于根据缓存层级公开缓存的提示 token。此更新是一个发布候选版本,表明它是在稳定版本发布前用于测试和反馈的预发布版本。该版本由 Cam Quilici 标记,并包含 Nick Hill 和 Yifan Qiao 的贡献。 AI

影响 vLLM 是一个用于高效 LLM 推理的开源库,此次更新提供了关于缓存提示 token 的新指标,可能有助于开发人员优化推理性能。

排序理由 这是一个开源库的发布候选版本,而不是重大的新模型或产品发布。

在 vLLM — Releases 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

vLLM 发布 v0.31.1rc0,新增缓存指标

本文如何被排名

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是一个开源库的发布候选版本,而不是重大的新模型或产品发布。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. vLLM — Releases TIER_1 English(EN) · cquil11 ·

    v0.31.1rc0: [指标] 按缓存层级公开缓存的提示令牌 (#56318)

    <p>Signed-off-by: Cam Quilici <a href="mailto:[email protected]">[email protected]</a><br /> Signed-off-by: Cam Quilici <a href="mailto:[email protected]">[email protected]</a><br /> Co-authored-by: Cam Quilici <a href="mailto:[email protected]">cameron@s…