PulseAugur
EN
LIVE 04:10:30

llama.cpp project releases pre-release updates with various fixes and optimizations

The llama.cpp project has released several pre-release updates, including b11514 which addresses Musa FWHT fixes and platform attestations for various operating systems. Other releases like b11512 and b11511 focus on model fixes and CUDA optimizations, respectively. Notably, b11509 includes a fix for CUDA CCCL version guards, and b11505 resolves issues with Vulkan TOP_K for infinite and NaN inputs. The updates also feature contributions from various developers and integrations with libraries like cpp-httplib. AI

IMPACT These updates improve the performance and stability of running LLMs on various platforms.

RANK_REASON The cluster consists of multiple pre-release updates to the llama.cpp project, which is a software tool for running large language models.

Read on llama.cpp — Releases →

AI-generated summary · Google Gemini · from 12 sources. How we write summaries →

llama.cpp project releases pre-release updates with various fixes and optimizations

How we ranked this

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster consists of multiple pre-release updates to the llama.cpp project, which is a software tool for running large language models.
Source corroboration
12 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

Full methodology in our editorial standards.

COVERAGE [12]

  1. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11514

    <details open=""> <p>Musa FWHT fix (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30167">#30167</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> </ul> <p><str…

  2. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11512

    <details open=""> <p>model : fix DFlash output head sharing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30111">#30111</a>)</p> <ul> <li>llama : fix DFlash output head sharing</li> </ul> <p>Assisted-by: Codex</p> <ul> <li>dflash : read tie…

  3. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11511

    <details open=""> <p>CUDA: fix MMQ out-of-bounds reads (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29953">#29953</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…

  4. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11510

    <details open=""> <p>CUDA : looped PAD kernel for more than 65535 rows or slices (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30147">#30147</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofo…

  5. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11509

    <details open=""> <p>CUDA: fix CCCL version guard breaking on major version rollover (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29453">#29453</a>)</p> <ul> <li>CUDA: fix CCCL version guard breaking on major version rollover</li> </ul> <p…

  6. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11507

    <details open=""> <p>llama: support MoE cache over multiple GPUs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30112">#30112</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://ll…

  7. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11505

    <details open=""> <p>vulkan : fix TOP_K for +inf/NaN inputs and k = 1 on negative values (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30107">#30107</a>)</p> <p>The bucket search in topk_nary_search.comp started from the range<br /> [0, 0xF…

  8. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11503

    <details open=""> <p>vendor : update cpp-httplib to 0.60.1 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30134">#30134</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </details> <p…

  9. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11501

    <details open=""> <p>sycl: FWHT optimizations (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29605">#29605</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> </…

  10. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11500

    <details open=""> <p>hex-mmadd: do not assume aligned read/write when bias-add is fused (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30133">#30133</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" re…

  11. llama.cpp — Releases TIER_1 (SO) · jeffbolznv ·

    b11504

    <p>vulkan: extend sparse FA support to coopmat2 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30003">#30003</a>)</p>

  12. llama.cpp — Releases TIER_1 (SO) · saady789 ·

    b11502

    <p>cuda : support arbitrary striding for unary ops on f16, f32, and bf16…</p>