PulseAugur
中
实时 20:28:18
English(EN) b11422: cuda: use the vector lightning indexer kernel on MUSA (#29990)

llama.cpp b11422 通过新的 MUSA indexer kernel 优化 CUDA/ROCm

llama.cpp 项目发布了 b11422 版本,其中包括对 CUDA 和 ROCm 架构的优化。具体来说,此次更新为 MUSA 引入了一个新的 vector lightning indexer kernel,解决了某些 MUSA 架构上的共享内存限制。通过分批处理,这一更改确保了具有 28 KB 静态共享内存上限的 MUSA 架构 21 和 22 可以有效地处理 indexer 查询。 AI

影响 在特定硬件配置上提升 AI 推理性能。

排序理由 这是对一个开源项目的软件更新,旨在优化特定硬件的性能,而不是一项新发布或重要的行业事件。

在 llama.cpp — Releases 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

llama.cpp b11422 通过新的 MUSA indexer kernel 优化 CUDA/ROCm

本文如何被排名

Signal score
5 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是对一个开源项目的软件更新,旨在优化特定硬件的性能,而不是一项新发布或重要的行业事件。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. llama.cpp — Releases TIER_1 English(EN) · ServeurpersoCom ·

    b11422: cuda: 在 MUSA 上使用 vector lightning indexer 内核 (#29990)

    <ul> <li>cuda: stage the lightning indexer queries in head passes for MUSA</li> </ul> <p>MUSA archs 21 and 22 cap static shared memory at 28 KB, and the tile<br /> kernel staged the queries of all four heads next to the key tile for<br /> 33 KB. The queries are now staged in pass…