PulseAugur
中
实时 04:38:04

llama.cpp项目发布预发布更新,包含多项修复和优化

llama.cpp项目已发布数个预发布更新,其中包括b11514,该更新解决了Musa FWHT修复和针对多个操作系统的平台证明问题。其他更新如b11512和b11511分别侧重于模型修复和CUDA优化。值得注意的是,b11509包含对CUDA CCCL版本保护的修复,而b11505解决了Vulkan TOP_K在处理无限和NaN输入时的问题。这些更新还包括了多位开发者的贡献以及与cpp-httplib等库的集成。 AI

影响 这些更新提高了在各种平台上运行LLM的性能和稳定性。

排序理由 该集群包含llama.cpp项目的多个预发布更新,llama.cpp是一个用于运行大型语言模型的软件工具。

在 llama.cpp — Releases 阅读 →

AI 生成摘要 · Google Gemini · 来自 12 个来源。 我们如何撰写摘要 →

llama.cpp项目发布预发布更新,包含多项修复和优化

本文如何被排名

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含llama.cpp项目的多个预发布更新,llama.cpp是一个用于运行大型语言模型的软件工具。
Source corroboration
12 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [12]

  1. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11514

    <details open=""> <p>Musa FWHT fix (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30167">#30167</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> </ul> <p><str…

  2. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11512

    <details open=""> <p>model : fix DFlash output head sharing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30111">#30111</a>)</p> <ul> <li>llama : fix DFlash output head sharing</li> </ul> <p>Assisted-by: Codex</p> <ul> <li>dflash : read tie…

  3. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11511

    <details open=""> <p>CUDA: fix MMQ out-of-bounds reads (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29953">#29953</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…

  4. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11510

    <details open=""> <p>CUDA : looped PAD kernel for more than 65535 rows or slices (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30147">#30147</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofo…

  5. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11509

    <details open=""> <p>CUDA: fix CCCL version guard breaking on major version rollover (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29453">#29453</a>)</p> <ul> <li>CUDA: fix CCCL version guard breaking on major version rollover</li> </ul> <p…

  6. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11507

    <details open=""> <p>llama: support MoE cache over multiple GPUs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30112">#30112</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://ll…

  7. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11505

    <details open=""> <p>vulkan : fix TOP_K for +inf/NaN inputs and k = 1 on negative values (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30107">#30107</a>)</p> <p>The bucket search in topk_nary_search.comp started from the range<br /> [0, 0xF…

  8. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11503

    <details open=""> <p>vendor : update cpp-httplib to 0.60.1 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30134">#30134</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </details> <p…

  9. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11501

    <details open=""> <p>sycl: FWHT optimizations (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29605">#29605</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> </…

  10. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11500

    <details open=""> <p>hex-mmadd: do not assume aligned read/write when bias-add is fused (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30133">#30133</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" re…

  11. llama.cpp — Releases TIER_1 (SO) · jeffbolznv ·

    b11504

    <p>vulkan: extend sparse FA support to coopmat2 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30003">#30003</a>)</p>

  12. llama.cpp — Releases TIER_1 (SO) · saady789 ·

    b11502

    <p>cuda : support arbitrary striding for unary ops on f16, f32, and bf16…</p>