PulseAugur
EN
LIVE 12:39:18

llama.cpp releases include server fixes, OpenVINO optimizations, and expanded platform support

The llama.cpp project has released several updates, including versions b11379, b11378, b11377, b11376, b11375, b11374, b11372, b11371, and b11370. These releases introduce a variety of improvements and fixes across different platforms and functionalities. Notable changes include server-side fixes for aborts, enhancements to chat parsing with JSON schema, optimizations for OpenVINO, and updates to CUDA and Metal backends. The project also continues to expand its support for different hardware architectures and operating systems, such as Snapdragon and openEuler. AI

IMPACT Ongoing development and optimization for local LLM inference tools.

RANK_REASON This is a series of software release notes for a specific project, not a major industry event.

Read on llama.cpp — Releases →

AI-generated summary · Google Gemini · from 73 sources. How we write summaries →

llama.cpp releases include server fixes, OpenVINO optimizations, and expanded platform support

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
This is a series of software release notes for a specific project, not a major industry event.
Source corroboration
73 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
7 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+17 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [73]

  1. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11399

    <details open=""> <p>CUDA: refactor swizzling code (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29612">#29612</a>)</p> <ul> <li> <p>CUDA: refactor swizzling code</p> </li> <li> <p>fix templates/loop bounds</p> </li> </ul> </details> <p><st…

  2. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11398

    <details open=""> <p>ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29806">#29806</a>)</p> <ul> <li> <p>ggml-cpu: vectorize BF16 K tails in tinyBLAS</p> </li> <li> <p>tests: Skip ti…

  3. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11397

    <details open=""> <p>cuda : move neu_padded to where it is used (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29940">#29940</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </detail…

  4. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11396

    <details open=""> <p>ci : windows llvm build requires ninja multi-config (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29959">#29959</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">ht…

  5. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11393

    <details open=""> <p>chat-peg-parser : clear current_tool when pending_tool_call is reset (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29942">#29942</a>)</p> <p>A TOOL_ID node that arrives after TOOL_CLOSE wrote through <code>current_tool<…

  6. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11392

    <details open=""> <p>ci : set default permissions (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29945">#29945</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li…

  7. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11391

    <details open=""> <p>cuda : move blocks_per_col to where it is used (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29939">#29939</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </de…

  8. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11390

    <details open=""> <p>CUDA: fix MMQ memory fault if n_expert &gt;&gt; n_ubatch (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29941">#29941</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollo…

  9. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11389

    <details open=""> <p>vulkan: fix rdna4 mat_vec tuning (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29934">#29934</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a>…

  10. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11388

    <details open=""> <p>imatrix: calculate activation-based statistics for new format (GGUF) imatrices (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/14891">#14891</a>)</p> <ul> <li>Use activations to calculate the stats</li> <li>Determine calc…

  11. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11387

    <details open=""> <p>spec : fix n-gram drafts rejected at temp &gt; 0 after truncation (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29924">#29924</a>)</p> <p>Co-authored-by: Pranesh Gonegandla <a href="mailto:[email protected]">pgoneg…

  12. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11386

    <details open=""> <p>common : prepare load_from_models_dir() for path conversion (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29674">#29674</a>)</p> <p>This is part of the fs::path modernization series.<br /> That was also the opportunity …

  13. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11385

    <details open=""> <p>server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29938">#29938</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">angt@huggingf…

  14. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11384

    <details open=""> <p>ci : pushing tag needs deploy key (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29937">#29937</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…

  15. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11382

    <details open=""> <p>webgpu: add f16 support to fill/set_rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29897">#29897</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama…

  16. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11381

    <details open=""> <p>mtmd : fix deprecated strdup warning on Windows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29863">#29863</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </d…

  17. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11380

    <details open=""> <p>vendor : update cpp-httplib to 0.59.0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29886">#29886</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.ap…

  18. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11379

    <details open=""> <p>server : fix laya abort by limiting n_batch to n_ubatch (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29903">#29903</a>)</p> <ul> <li>server : fix laya abort by limiting n_batch to n_ubatch</li> </ul> <p>Fixes <a class=…

  19. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11378

    <details open=""> <p>common : add common_is_tty() helper and fix deprecated warnings on Windows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29860">#29860</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">angt…

  20. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11377

    <details open=""> <p>chat : honor json_schema in Ling 3.0 parser (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29813">#29813</a>)</p> <ul> <li>chat : honor json_schema in Ling 3.0 parser</li> </ul> <p>Ling 3.0 only built a grammar for tool …

  21. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11376

    <details open=""> <p>ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29904">#29904</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofol…

  22. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11375

    <details open=""> <p>graph: gather the recurrent states once so the reserve covers every split (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29856">#29856</a>)</p> <p>build_rs gathered the extra states (n_rs - n_seqs rows) with their own<br…

  23. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11374

    <details open=""> <p>ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29852">#29852</a>)</p> <ul> <li>ggml-openvino : Qwen3.5 MoE perf (<a class="issu…

  24. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11372

    <details open=""> <p>qwen4exp : halve the indexer score memory (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29825">#29825</a>)</p> <ul> <li>qwen4exp : halve the indexer score memory</li> </ul> <p>The indexer scored all heads in one product…

  25. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11371

    <details open=""> <p>model: add support for clef decision model (text-only) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29831">#29831</a>)</p> <ul> <li> <p>init support for clef (text only)</p> </li> <li> <p>more static graph</p> </li> <l…

  26. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11370

    <details open=""> <p>CUDA: fuse shared experts into MMVQ (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29184">#29184</a>)</p> <ul> <li> <p>CUDA: fuse shared experts into MMVQ</p> </li> <li> <p>check if buffer is null</p> </li> <li> <p>move …

  27. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11368

    <details open=""> <p>spec : add probabilistic sampling for simple draft and MTP (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27694">#27694</a>)</p> <ul> <li> <p>Make the drafter probabilistic and the target verify by rejection sampling</p>…

  28. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11366

    <details open=""> <p>ggml-quants : avoid invalid rounding in qkx3 scale search (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29817">#29817</a>)</p> <ul> <li>ggml-quants : avoid invalid rounding in qkx3 scale search</li> </ul> <p>The imatrix…

  29. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11365

    <details open=""> <p>ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27096">#27096</a>)</p> <ul> <li>ggml-cpu : fix soft_max_back wrong output when dst aliases src1</li> </ul> <p…

  30. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11364

    <details open=""> <p>model: support nimble decision model (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29844">#29844</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app…

  31. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11362

    <details open=""> <p>metal : add tensor API flash attention kernel for F16 KV (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29570">#29570</a>)</p> <ul> <li> <p>metal : add tensor API flash attention kernel for F16 KV</p> </li> <li> <p>cont …

  32. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11361

    <details open=""> <p>llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29818">#29818</a>)</p> <ul> <li> <p>init conversion</p> </li> <li> <p>convert: ok</p> </li> <…

  33. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11355

    <details open=""> <p>vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28531">#28531</a>)</p> <p>Assisted-by: Claude Opus</p> </details> <p><strong>Website:</strong></p> …

  34. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11352

    <details open=""> <p>qwen4exp : optimize mask constructions (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29824">#29824</a>)</p> <ul> <li> <p>qwen4exp : optimize mask constructions</p> </li> <li> <p>cont : apply the same change for GLM5-nex…

  35. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11351

    <details open=""> <p>ggml : add <code>alloc_buffer_n</code> to buffer type interface (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23671">#23671</a>)</p> <ul> <li>ggml : add <code>alloc_buffer_n</code> to buffer type interface</li> </ul> <p…

  36. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11349

    <details open=""> <p>vulkan: add logging to pipeline compile issues (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29794">#29794</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:/…

  37. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11346

    <details open=""> <p>qwen4exp: fix tests (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29819">#29819</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> </ul> <…

  38. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11345

    <details open=""> <p>hexagon: add q2_k and q3_k quant type support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29717">#29717</a>)</p> <ul> <li> <p>hexagon: add q2_k and q3_k quant type support</p> </li> <li> <p>hex-qk: consistent allocati…

  39. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11344

    <details open=""> <p>CUDA: fix 2 broken Volta FA cases (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29803">#29803</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…

  40. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11342

    <details open=""> <p>common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29816">#29816</a>)</p> <p>See <a href="https://gcc.gnu.org/bugzilla/show_bug.cgi?id=101510" rel="nofo…

  41. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11339

    <details open=""> <p>llama : clamp kpool re-pool bound to existing pools (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29805">#29805</a>)</p> <ul> <li> <p>tests : simplify function signature</p> </li> <li> <p>llama : clamp kpool re-pool bou…

  42. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11338

    <details open=""> <p>hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29685">#29685</a>)</p> <ul> <li> <p>hexagon: shared strided DMA copy for CPY and CONCAT, any-dim …

  43. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11337

    <details open=""> <p>server: return HTTP 400 for invalid embedding requests (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29060">#29060</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow"…

  44. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11335

    <details open=""> <p>cuda : route sm70 to the Turing MMVQ nwarps table (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29753">#29753</a>)</p> <ul> <li>cuda : route sm70 to the Turing MMVQ nwarps table</li> </ul> <p>Volta (sm_70) has no MMVQ p…

  45. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11334

    <details open=""> <p>metal : release temporary private transfer buffers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29777">#29777</a>)</p> <ul> <li>metal : release temporary private transfer buffers</li> </ul> <p>Assisted-by: OpenAI Codex…

  46. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11333

    <details open=""> <p>webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29358">#29358</a> (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29358">#…

  47. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11332

    <details open=""> <p>llama : fix invalid assert in recurrent memory (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29799">#29799</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:/…

  48. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11331

    <details open=""> <p>CUDA: Handle compute type for NVFP4 on cublass path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29173">#29173</a>)</p> <ul> <li>CUDA: Handle compute type for NVFP4 on cublass path</li> </ul> <p>Signed-off-by: ynankani…

  49. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11330

    <details open=""> <p>Qwen4Exp: add MTP (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29761">#29761</a>)</p> <ul> <li> <p>Qwen4Exp: add MTP</p> </li> <li> <p>remove has_state member, check via ctx_bufs being non-empty</p> </li> <li> <p>consi…

  50. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11327

    <details open=""> <p>mtmd: cap max_image to ubatch for non_causal models (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29773">#29773</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">ht…

  51. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11326

    <details open=""> <p>meta: clear inactive AllReduce shards with FILL, not SCALE (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29793">#29793</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofol…

  52. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11325

    <details open=""> <p>jinja : skip copying loop scope unless a loop filter needs it (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29776">#29776</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="no…

  53. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11324

    <details open=""> <p>llama-mmap : avoid a second full-size copy of each tensor with direct-io (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29749">#29749</a>)</p> <p>Assisted-by: Claude</p> <p>Co-authored-by: Pranesh Gonegandla <a href="mai…

  54. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11323

    <details open=""> <p>HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29572">#29572</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llam…

  55. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11322

    <details open=""> <p>hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29785">#29785</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https:…

  56. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11321

    <details open=""> <p>BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29640">#29640</a>)</p> <ul> <li> <p>BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS</p> </li> <…

  57. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11320

    <details open=""> <p>common : add LLM-jp-4.1 Harmony dialect handler (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29681">#29681</a>)</p> <p>LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space<br /> after every special tok…

  58. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11319

    <details open=""> <p>opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29698">#29698</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" …

  59. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11318

    <details open=""> <p>vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29734">#29734</a>)</p> <ul> <li>vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3</li> </ul> <p>The original toke…

  60. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11317

    <details open=""> <p>llama-bench : fix verbosity filter to show GGML_LOG_ERROR (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28229">#28229</a>)</p> <ul> <li> <p>bench : fix verbosity filter to show GGML_LOG_ERROR (<a class="issue-link js-is…

  61. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11316

    <details open=""> <p>metal : use bf16 math for mxfp4 mul-mat (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29770">#29770</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.…

  62. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11313

    <details open=""> <p>model : re-enable -sm tensor for qwen4exp (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28569">#28569</a>)</p> <p><a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27941">#27941</a> di…

  63. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11312

    <details open=""> <p>webgpu: fix SSM_SCAN binding aliasing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29750">#29750</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.ap…

  64. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11311

    <details open=""> <p>ggml-opencl : replace alloca() with std::vector (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29765">#29765</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </d…

  65. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11310

    <details open=""> <p>cuda: guard the iq4_nl dequantize row kernel against short rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29683">#29683</a>)</p> <p>dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than…

  66. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11309

    <details open=""> <p>Hexagon: optimize ALLREDUCE with support for safe scatter mode (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29757">#29757</a>)</p> <ul> <li> <p>hex-allreduce: add support for safe scatter mode</p> </li> <li> <p>hex-all…

  67. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11308

    <details open=""> <p>args: fix cli download mmproj arg (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28977">#28977</a>)</p> <ul> <li> <p>tests: add tests for cli download arg parsing</p> </li> <li> <p>args: fix cli download mmproj arg</p> <…

  68. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11307

    <details open=""> <p>llama : preserve original batch order for speculative decoding layer inputs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29019">#29019</a>)</p> <ul> <li>llama: preserve original batch order for layer inputs</li> </ul> …

  69. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11306

    <details open=""> <p>test-llama-archs : toggle causal_attn to catch graph shape changes (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29724">#29724</a>)</p> <p>After the device decode, flip causal_attn off, decode n_ubatch/2 then<br /> n_ub…

  70. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11304

    <details open=""> <p>llama: properly handle KV on training (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28520">#28520</a>)</p> <ul> <li> <p>llama: properly handle KV on training</p> </li> <li> <p>improve</p> </li> </ul> </details> <p><stro…

  71. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11303

    <details open=""> <p>batch: migrate the rest of examples to llama_batch_ext (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29601">#29601</a>)</p> <ul> <li> <p>migrate the rest</p> </li> <li> <p>test-thread-safety</p> </li> <li> <p>rm common_…

  72. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11302

    <details open=""> <p>glm5-next: give dead indexer slots unique scatter rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29745">#29745</a>)</p> <p>The sparse indexer mask is built with a set_rows scatter. Padded pools,<br /> absent sequenc…

  73. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11301

    <details open=""> <p>ggml/gguf : fix integer overflow (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29384">#29384</a>)</p> <ul> <li> <p>ggml: fix integer overflow guard for zero-element tensors</p> </li> <li> <p>ggml: validate number of ele…