PulseAugur
中
实时 02:45:42
English(EN) llama-bench skipped FA on capable GPUs — b9437 corrects it

llama-bench 针对闪存注意力和 GPU 层数进行了默认值更正

最近为 llama-bench 工具发布的 b9437 版本更正了与闪存注意力和 GPU 层数相关的默认设置。此前,该工具即使在兼容硬件上也将闪存注意力硬编码为关闭,并为 GPU 层数使用了旧的哨兵值。此次更新现在将闪存注意力默认设置为在 सक्षम 硬件(CUDA、Metal、Vulkan)上自动激活,并将 GPU 层数设置为 -1,与其他 llama.cpp 工具(如 llama-server 和 llama-cli)保持一致。此更改确保了使用最新默认值运行的基准测试能够准确反映在支持的 GPU 上使用闪存注意力的情况。 AI

影响 确保在兼容硬件上对闪存注意力的准确基准测试,提高 llama.cpp 性能指标的可靠性。

排序理由 这是对特定工具默认设置的修复,而不是新的模型发布或重大的行业事件。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

llama-bench 针对闪存注意力和 GPU 层数进行了默认值更正

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是对特定工具默认设置的修复,而不是新的模型发布或重大的行业事件。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
103 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Creeta ·

    llama-bench 在强大 GPU 上跳过了 FA — b9437 进行了修正

    <h2> What flipped in b9437 </h2> <p>Build <a href="https://github.com/ggml-org/llama.cpp/releases" rel="noopener noreferrer">b9437</a>, published on May 30, 2026 at 20:56 UTC , ships two targeted default-value corrections to <code>llama-bench</code>. Flash attention (<code>-fa</c…