llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 20:52
<details open=""> <p>cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29722">#29722</a>)</p> <ul> <li>cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast</li> </ul…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 20:22
<details open=""> <p>mimo : support dflash (convert + feature extraction) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29650">#29650</a>)</p> <ul> <li> <p>convert : update to support dflash</p> </li> <li> <p>cont : fix</p> </li> </ul> <p>C…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 19:40
<details open=""> <p>jinja : support coerced array attributes (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29574">#29574</a>)</p> <ul> <li> <p>support coerced array attributes</p> </li> <li> <p>add tests</p> </li> </ul> </details> <p><stro…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 18:53
<details open=""> <p>ci : fix Models Backend Check by shortening the hrm_text fixture (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29744">#29744</a>)</p> <p>The fixture recycles its two blocks over 8 cache slots, so the fp16<br /> error bu…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 18:17
<details open=""> <p>llama: llama_prefetch_rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29599">#29599</a>)</p> <ul> <li> <p>llama: llama_prefetch_rows</p> </li> <li> <p>llama: support row prefetch on Windows</p> </li> </ul> <p>Apply t…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 17:38
<details open=""> <p>ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29675">#29675</a>)</p> <ul> <li> <p>ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)</p> </li> <li> …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 16:21
<details open=""> <p>cpu: accept BF16 in src1 of mul_mat (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28937">#28937</a>)</p> <ul> <li>cpu: accept BF16 in src1 of mul_mat</li> </ul> <p>ggml_conv_1d_dw builds its im2col in F32 when the kerne…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 12:39
<details open=""> <p>openvino: serve GET_ROWS on a weight view from the base Constant (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28381">#28381</a>)</p> <ul> <li>openvino: serve GET_ROWS on a weight view from the base Constant</li> </ul> …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 11:59
<details open=""> <p>musa : define <strong>CUDA_ARCH</strong> for device passes (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29508">#29508</a>)</p> <p>The MUSA vendor header never defined <strong>CUDA_ARCH</strong>, so every architecture<b…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 11:30
<details open=""> <p>SYCL: reduce tensor allreduce sync with pinned host buffers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29604">#29604</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofo…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 11:01
<details open=""> <p>add GLM-5.3-Flash (GLM5-Next) support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27773">#27773</a>)</p> <ul> <li> <p>Rebase GLM-Next support onto master, and migrate to llama-memory-hybrid-idx</p> </li> <li> <p>Add i…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 10:35
<details open=""> <p>vendor: update BoringSSL to 0.20260929.0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29669">#29669</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 09:01
<details open=""> <p>Hexagon f16 activation ops (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29209">#29209</a>)</p> <ul> <li>hexagon: add F16 support for activation ops (SILU/GELU/GELU_QUICK/GEGLU/SWIGLU)</li> </ul> <p>Widens ggml_hexagon_…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 08:33
<details open=""> <p>gguf : reject tensor size that wraps after padding (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26979">#26979</a>)</p> <p>GGML_PAD(nbytes, alignment) wraps to 0 when nbytes is within<br /> (alignment - 1) of SIZE_MAX, …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 07:58
<details open=""> <p>ggml : check row bounds in get_rows_back (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29575">#29575</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected] ">[email protected] </a></p> </details>…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 07:25
<details open=""> <p>model : support classifier_pooling for rerankers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29627">#29627</a>)</p> <ul> <li>model : support classifier_pooling for ModernBERT rerankers</li> </ul> <p>Assisted-by: Claud…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 06:14
<details open=""> <p>hexagon: optimize concat op (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29673">#29673</a>)</p> <ul> <li>hex-concat: reduce pkts in gather/transpose hot loop</li> </ul> <p>gather directly into dst buffer, use special i…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 03:56
<details open=""> <p>ggml-cuda: HIP: optimize packed byte subtraction (<code>__vsubss4</code> -> <code>__vsub4</code>) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29478">#29478</a>)</p> <ul> <li> <p>ggml-cuda: HIP: optimize non-saturat…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 01:21
<details open=""> <p>ci: add zdnn backend build but not test (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29541">#29541</a>)</p> <ul> <li>ci: add zdnn backend build but not test</li> </ul> <p>Signed-off-by: Aaron Teo <a href="mailto:aaron.…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 00:49
<details open=""> <p>opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29555">#29555</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="no…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-30 00:19
<details open=""> <p>vulkan: Tune GDN kernel, fix Intel performance (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29476">#29476</a>)</p> <ul> <li> <p>vulkan: tune GDN shader</p> </li> <li> <p>tune for Intel</p> </li> </ul> </details> <p><st…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 23:54
<details open=""> <p>vulkan : Load F32 A matrix 2 at a time when its 2-aligned (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29254">#29254</a>)</p> <p>It turns out Intel doesn't particularly like loading F32s one at a<br /> time and we alre…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 23:27
<details open=""> <p>vulkan: MOE aware mat_mul_id tile selection (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29182">#29182</a>)</p> <p>mut_mul_id selected its matmul tile with total token count.<br /> For MoE dispatch grid the true N per …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 23:00
<details open=""> <p>ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29504">#29504</a>)</p> <ul> <li> <p>fix c++ odr by properly using GGML_COMMON_DECL_CPP</p> </li> <li> <p>using actu…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 22:35
<details open=""> <p>vocab : keep NORMAL in PLaMo-2 and PLaMo-3 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29580">#29580</a>)</p> <ul> <li>vocab : keep NORMAL in PLaMo-2 and PLaMo-3</li> </ul> <p>The PLaMo-2 and PLaMo-3 vocabularies mark…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 21:53
<details open=""> <p>ggml : accumulate f16 dot products in f32 on AVX512-FP16 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29545">#29545</a>)</p> <p>Supersedes <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 21:22
<details open=""> <p>ggml : require input tensors to be GGML_OP_NONE (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29647">#29647</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 19:59
<details open=""> <p>hexagon: add FP32 GELU_ERF and GEGLU_ERF support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29631">#29631</a>)</p> <ul> <li> <p>hexagon: add FP32 GELU_ERF and GEGLU_ERF support</p> </li> <li> <p>hex-erf: reduce regis…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 17:17
<details open=""> <p>common : stop accepting draft tokens at EOG (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29638">#29638</a>)</p> <ul> <li> <p>common : stop accepting draft tokens at EOG</p> </li> <li> <p>cont : remove the test</p> </li…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 16:33
<details open=""> <p>server : remove the built-in UI's service worker when the UI is not served (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29565">#29565</a>)</p> <p>With --path or --no-ui, /sw.js returned 404, and a 404 does not remove a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 16:02
<details open=""> <p>common : use fs::path for config dir (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29649">#29649</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected] ">[email protected] </a></p> </details> <p>…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 13:10
<details open=""> <p>tests : adjust server string regex to also match m2 utlra results (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29648">#29648</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 12:31
<details open=""> <p>common : add fs_write_atomic() (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29642">#29642</a>)</p> <ul> <li>Check for buffered write errors when closing downloaded files.</li> <li>Use UTF-8 paths when writing ETag file…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 11:22
<details open=""> <p>ggml : collect all input tensors into graph_inputs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29634">#29634</a>)</p> <p>graph_inputs was populated while splitting the graph, so it only<br /> contained the inputs that…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 09:40
<details open=""> <p>llama : fix init in several tools/examples (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29632">#29632</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://lla…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 07:31
<details open=""> <p>vulkan : reuse descriptor sets when bindings are constant (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29280">#29280</a>)</p> <ul> <li> <p>vulkan : reuse descriptor sets when bindings are constant</p> </li> <li> <p>vul…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 06:49
<details open=""> <p>chat : fix Muse Glimmer ignoring response_format json_schema with --jinja (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29615">#29615</a>)</p> <ul> <li>chat : fix Muse Glimmer ignoring response_format json_schema with -…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-29 06:16
<details open=""> <p>common : use fs::path for cache dirs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29595">#29595</a>)</p> <ul> <li>Avoid useless string conversions on Windows.</li> <li>No need for BSD or emscripten special cases.</li> …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-28 22:30
<details open=""> <p>server : support typed content (vision/audio/video) input for /v1/embeddings endpoint (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29556">#29556</a>)</p> <ul> <li>server : support multimodal input for /v1/embeddings (Q…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-28 21:31
<details open=""> <p>models: pad on the left with ggml_pad_ext (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29567">#29567</a>)</p> <ul> <li>models: pad on the left with ggml_pad_ext</li> </ul> <p>The Parakeet, LFM2-Audio, Granite Speech an…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-28 20:55
<details open=""> <p>ggml-openvino: mark unaligned batch-stride views unsupported (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29603">#29603</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nof…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-28 18:35
<details open=""> <p>batch: migrate speculative, mtmd and server to batch_ext (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29385">#29385</a>)</p> <ul> <li> <p>adapt common</p> </li> <li> <p>add common_batch</p> </li> <li> <p>wip</p> </li> …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-28 16:49
<details open=""> <p>common : fix HF cache paths on Windows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29475">#29475</a>)</p> <p>Supersedes <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29158">#2915…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-28 16:16
<details open=""> <p>webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29471">#29471</a>)</p> <ul> <li> <p>Fix: Handle unaligned writes in ggml_backend_webgpu_buffer_set_t…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-28 15:37
<details open=""> <p>tests : refactor test-recurrent-state-rollback (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29426">#29426</a>)</p> <ul> <li>tests : use llama_context_ptr in test-recurrent-state-rollback</li> </ul> <p>Replace raw llama…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-28 14:25
<details open=""> <p>ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29423">#29423</a>)</p> <ul> <li> <p>ggml-cpu: enable tiled flash attention for non-vector-mul…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-28 12:45
<details open=""> <p>HIP: fix template skip for DKQ > 256 mfma kernels (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29559">#29559</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">h…