llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 22:56
<details open=""> <p>CUDA: refactor swizzling code (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29612">#29612</a>)</p> <ul> <li> <p>CUDA: refactor swizzling code</p> </li> <li> <p>fix templates/loop bounds</p> </li> </ul> </details> <p><st…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 20:07
<details open=""> <p>ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29806">#29806</a>)</p> <ul> <li> <p>ggml-cpu: vectorize BF16 K tails in tinyBLAS</p> </li> <li> <p>tests: Skip ti…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 19:50
<details open=""> <p>cuda : move neu_padded to where it is used (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29940">#29940</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected] ">[email protected] </a></p> </detail…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 17:49
<details open=""> <p>ci : windows llvm build requires ninja multi-config (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29959">#29959</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">ht…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 16:25
<details open=""> <p>chat-peg-parser : clear current_tool when pending_tool_call is reset (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29942">#29942</a>)</p> <p>A TOOL_ID node that arrives after TOOL_CLOSE wrote through <code>current_tool<…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 14:43
<details open=""> <p>ci : set default permissions (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29945">#29945</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 14:12
<details open=""> <p>cuda : move blocks_per_col to where it is used (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29939">#29939</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected] ">[email protected] </a></p> </de…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 12:56
<details open=""> <p>CUDA: fix MMQ memory fault if n_expert >> n_ubatch (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29941">#29941</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollo…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 12:34
<details open=""> <p>vulkan: fix rdna4 mat_vec tuning (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29934">#29934</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a>…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 10:42
<details open=""> <p>imatrix: calculate activation-based statistics for new format (GGUF) imatrices (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/14891">#14891</a>)</p> <ul> <li>Use activations to calculate the stats</li> <li>Determine calc…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 10:22
<details open=""> <p>spec : fix n-gram drafts rejected at temp > 0 after truncation (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29924">#29924</a>)</p> <p>Co-authored-by: Pranesh Gonegandla <a href="mailto:[email protected] ">pgoneg…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 10:03
<details open=""> <p>common : prepare load_from_models_dir() for path conversion (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29674">#29674</a>)</p> <p>This is part of the fs::path modernization series.<br /> That was also the opportunity …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 09:37
<details open=""> <p>server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29938">#29938</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected] ">angt@huggingf…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 08:06
<details open=""> <p>ci : pushing tag needs deploy key (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29937">#29937</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-04 00:37
<details open=""> <p>webgpu: add f16 support to fill/set_rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29897">#29897</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 21:35
<details open=""> <p>mtmd : fix deprecated strdup warning on Windows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29863">#29863</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected] ">[email protected] </a></p> </d…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 21:09
<details open=""> <p>vendor : update cpp-httplib to 0.59.0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29886">#29886</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.ap…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 15:54
<details open=""> <p>server : fix laya abort by limiting n_batch to n_ubatch (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29903">#29903</a>)</p> <ul> <li>server : fix laya abort by limiting n_batch to n_ubatch</li> </ul> <p>Fixes <a class=…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 15:02
<details open=""> <p>common : add common_is_tty() helper and fix deprecated warnings on Windows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29860">#29860</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected] ">angt…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 14:27
<details open=""> <p>chat : honor json_schema in Ling 3.0 parser (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29813">#29813</a>)</p> <ul> <li>chat : honor json_schema in Ling 3.0 parser</li> </ul> <p>Ling 3.0 only built a grammar for tool …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 13:27
<details open=""> <p>ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29904">#29904</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofol…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 12:40
<details open=""> <p>graph: gather the recurrent states once so the reserve covers every split (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29856">#29856</a>)</p> <p>build_rs gathered the extra states (n_rs - n_seqs rows) with their own<br…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 11:54
<details open=""> <p>ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29852">#29852</a>)</p> <ul> <li>ggml-openvino : Qwen3.5 MoE perf (<a class="issu…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 11:28
<details open=""> <p>qwen4exp : halve the indexer score memory (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29825">#29825</a>)</p> <ul> <li>qwen4exp : halve the indexer score memory</li> </ul> <p>The indexer scored all heads in one product…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 08:36
<details open=""> <p>model: add support for clef decision model (text-only) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29831">#29831</a>)</p> <ul> <li> <p>init support for clef (text only)</p> </li> <li> <p>more static graph</p> </li> <l…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 06:27
<details open=""> <p>CUDA: fuse shared experts into MMVQ (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29184">#29184</a>)</p> <ul> <li> <p>CUDA: fuse shared experts into MMVQ</p> </li> <li> <p>check if buffer is null</p> </li> <li> <p>move …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 02:38
<details open=""> <p>spec : add probabilistic sampling for simple draft and MTP (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27694">#27694</a>)</p> <ul> <li> <p>Make the drafter probabilistic and the target verify by rejection sampling</p>…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 01:58
<details open=""> <p>ggml-quants : avoid invalid rounding in qkx3 scale search (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29817">#29817</a>)</p> <ul> <li>ggml-quants : avoid invalid rounding in qkx3 scale search</li> </ul> <p>The imatrix…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 01:32
<details open=""> <p>ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27096">#27096</a>)</p> <ul> <li>ggml-cpu : fix soft_max_back wrong output when dst aliases src1</li> </ul> <p…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 01:06
<details open=""> <p>model: support nimble decision model (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29844">#29844</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-03 00:28
<details open=""> <p>metal : add tensor API flash attention kernel for F16 KV (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29570">#29570</a>)</p> <ul> <li> <p>metal : add tensor API flash attention kernel for F16 KV</p> </li> <li> <p>cont …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 23:57
<details open=""> <p>llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29818">#29818</a>)</p> <ul> <li> <p>init conversion</p> </li> <li> <p>convert: ok</p> </li> <…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 23:18
<details open=""> <p>vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28531">#28531</a>)</p> <p>Assisted-by: Claude Opus</p> </details> <p><strong>Website:</strong></p> …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 22:23
<details open=""> <p>qwen4exp : optimize mask constructions (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29824">#29824</a>)</p> <ul> <li> <p>qwen4exp : optimize mask constructions</p> </li> <li> <p>cont : apply the same change for GLM5-nex…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 18:55
<details open=""> <p>ggml : add <code>alloc_buffer_n</code> to buffer type interface (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23671">#23671</a>)</p> <ul> <li>ggml : add <code>alloc_buffer_n</code> to buffer type interface</li> </ul> <p…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 15:43
<details open=""> <p>vulkan: add logging to pipeline compile issues (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29794">#29794</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:/…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 13:52
<details open=""> <p>qwen4exp: fix tests (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29819">#29819</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> </ul> <…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 13:11
<details open=""> <p>hexagon: add q2_k and q3_k quant type support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29717">#29717</a>)</p> <ul> <li> <p>hexagon: add q2_k and q3_k quant type support</p> </li> <li> <p>hex-qk: consistent allocati…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 08:56
<details open=""> <p>CUDA: fix 2 broken Volta FA cases (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29803">#29803</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 05:09
<details open=""> <p>common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29816">#29816</a>)</p> <p>See <a href="https://gcc.gnu.org/bugzilla/show_bug.cgi?id=101510" rel="nofo…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 04:32
<details open=""> <p>llama : clamp kpool re-pool bound to existing pools (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29805">#29805</a>)</p> <ul> <li> <p>tests : simplify function signature</p> </li> <li> <p>llama : clamp kpool re-pool bou…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 04:07
<details open=""> <p>hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29685">#29685</a>)</p> <ul> <li> <p>hexagon: shared strided DMA copy for CPY and CONCAT, any-dim …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 03:37
<details open=""> <p>server: return HTTP 400 for invalid embedding requests (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29060">#29060</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow"…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 03:10
<details open=""> <p>cuda : route sm70 to the Turing MMVQ nwarps table (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29753">#29753</a>)</p> <ul> <li>cuda : route sm70 to the Turing MMVQ nwarps table</li> </ul> <p>Volta (sm_70) has no MMVQ p…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 02:37
<details open=""> <p>metal : release temporary private transfer buffers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29777">#29777</a>)</p> <ul> <li>metal : release temporary private transfer buffers</li> </ul> <p>Assisted-by: OpenAI Codex…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 01:57
<details open=""> <p>webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29358">#29358</a> (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29358">#…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 01:24
<details open=""> <p>llama : fix invalid assert in recurrent memory (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29799">#29799</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:/…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 00:51
<details open=""> <p>CUDA: Handle compute type for NVFP4 on cublass path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29173">#29173</a>)</p> <ul> <li>CUDA: Handle compute type for NVFP4 on cublass path</li> </ul> <p>Signed-off-by: ynankani…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-02 00:21
<details open=""> <p>Qwen4Exp: add MTP (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29761">#29761</a>)</p> <ul> <li> <p>Qwen4Exp: add MTP</p> </li> <li> <p>remove has_state member, check via ctx_bufs being non-empty</p> </li> <li> <p>consi…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 20:59
<details open=""> <p>mtmd: cap max_image to ubatch for non_causal models (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29773">#29773</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">ht…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 20:33
<details open=""> <p>meta: clear inactive AllReduce shards with FILL, not SCALE (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29793">#29793</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofol…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 20:02
<details open=""> <p>jinja : skip copying loop scope unless a loop filter needs it (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29776">#29776</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="no…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 19:28
<details open=""> <p>llama-mmap : avoid a second full-size copy of each tensor with direct-io (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29749">#29749</a>)</p> <p>Assisted-by: Claude</p> <p>Co-authored-by: Pranesh Gonegandla <a href="mai…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 18:48
<details open=""> <p>HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29572">#29572</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llam…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 18:03
<details open=""> <p>hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29785">#29785</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https:…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 14:13
<details open=""> <p>BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29640">#29640</a>)</p> <ul> <li> <p>BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS</p> </li> <…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 11:23
<details open=""> <p>common : add LLM-jp-4.1 Harmony dialect handler (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29681">#29681</a>)</p> <p>LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space<br /> after every special tok…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 10:48
<details open=""> <p>opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29698">#29698</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 10:20
<details open=""> <p>vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29734">#29734</a>)</p> <ul> <li>vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3</li> </ul> <p>The original toke…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 09:36
<details open=""> <p>llama-bench : fix verbosity filter to show GGML_LOG_ERROR (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28229">#28229</a>)</p> <ul> <li> <p>bench : fix verbosity filter to show GGML_LOG_ERROR (<a class="issue-link js-is…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 09:01
<details open=""> <p>metal : use bf16 math for mxfp4 mul-mat (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29770">#29770</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 08:08
<details open=""> <p>model : re-enable -sm tensor for qwen4exp (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28569">#28569</a>)</p> <p><a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27941">#27941</a> di…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 03:11
<details open=""> <p>webgpu: fix SSM_SCAN binding aliasing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29750">#29750</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.ap…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 02:49
<details open=""> <p>ggml-opencl : replace alloca() with std::vector (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29765">#29765</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected] ">[email protected] </a></p> </d…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 02:09
<details open=""> <p>cuda: guard the iq4_nl dequantize row kernel against short rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29683">#29683</a>)</p> <p>dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 01:38
<details open=""> <p>Hexagon: optimize ALLREDUCE with support for safe scatter mode (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29757">#29757</a>)</p> <ul> <li> <p>hex-allreduce: add support for safe scatter mode</p> </li> <li> <p>hex-all…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 01:09
<details open=""> <p>args: fix cli download mmproj arg (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28977">#28977</a>)</p> <ul> <li> <p>tests: add tests for cli download arg parsing</p> </li> <li> <p>args: fix cli download mmproj arg</p> <…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-10-01 00:37
<details open=""> <p>llama : preserve original batch order for speculative decoding layer inputs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29019">#29019</a>)</p> <ul> <li>llama: preserve original batch order for layer inputs</li> </ul> …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 23:36
<details open=""> <p>test-llama-archs : toggle causal_attn to catch graph shape changes (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29724">#29724</a>)</p> <p>After the device decode, flip causal_attn off, decode n_ubatch/2 then<br /> n_ub…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 23:08
<details open=""> <p>llama: properly handle KV on training (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28520">#28520</a>)</p> <ul> <li> <p>llama: properly handle KV on training</p> </li> <li> <p>improve</p> </li> </ul> </details> <p><stro…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 22:40
<details open=""> <p>batch: migrate the rest of examples to llama_batch_ext (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29601">#29601</a>)</p> <ul> <li> <p>migrate the rest</p> </li> <li> <p>test-thread-safety</p> </li> <li> <p>rm common_…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 21:59
<details open=""> <p>glm5-next: give dead indexer slots unique scatter rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29745">#29745</a>)</p> <p>The sparse indexer mask is built with a set_rows scatter. Padded pools,<br /> absent sequenc…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 21:25
<details open=""> <p>ggml/gguf : fix integer overflow (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29384">#29384</a>)</p> <ul> <li> <p>ggml: fix integer overflow guard for zero-element tensors</p> </li> <li> <p>ggml: validate number of ele…