PulseAugur
中
实时 19:52:24
English(EN) b11513: CUDA: improve top-k algorithm selection (#28713)

llama.cpp 优化 CUDA top-k 算法以实现显著加速

llama.cpp 项目发布了一个更新 (b11513),显著优化了 top-k 算法的 CUDA 实现。此更新用更高效的 grid-over-rows radix select 替换了 per-row DeviceTopKKernel,适用于行数较多的情况,从而大幅缩短了处理时间。例如,在 qwen4exp 模型上,使用 34,816 个 token 时,top-k 操作速度从超过 5.7 秒提高到大约 941 毫秒。该版本还根据形状优化了最佳 top-k 实现的选择,根据行长度和配置使用 bitonic、radix select 或 DeviceTopK/CUB argsort。 AI

影响 通过 llama.cpp 提高了在支持 CUDA 的硬件上运行模型的推理性能。

排序理由 这是开源项目的一个软件更新,它提高了特定算法在某些硬件上的性能,而不是前沿发布或重大的行业事件。

在 llama.cpp — Releases 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

llama.cpp 优化 CUDA top-k 算法以实现显著加速

本文如何被排名

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是开源项目的一个软件更新,它提高了特定算法在某些硬件上的性能,而不是前沿发布或重大的行业事件。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. llama.cpp — Releases TIER_1 English(EN) · praneshgo ·

    b11513: CUDA:改进 top-k 算法选择 (#28713)

    <ul> <li>CUDA: radix top-k for large row counts</li> </ul> <p>Replaces CUB's per-row DeviceTopKKernel with a grid-over-rows radix select,<br /> gated on GGML_CUDA_TOPK_RADIX_MIN_ROWS. On qwen4exp at 34,816 tokens this cuts<br /> top-k from 1,671,253 launches / 5,761.8 ms to 2,329…