llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 17:32
<details open=""> <p>vulkan: use spec constant for matrix matrix multiplication A-type (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25773">#25773</a>)</p> <ul> <li>vulkan: use spec constant for mul mat type_a</li> </ul> <p>vulkan: use map …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 15:34
<details open=""> <p>vulkan: Convert FILL to distribute workgroups in 2D to avoid exceeding maxComputeWorkGroupCount (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28592">#28592</a>)</p> <ul> <li>divide workload to 2D</li> </ul> <p>This is t…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 13:25
<details open=""> <p>llama : use int32_t for llama_sampler_chain_n return type (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28631">#28631</a>)</p> <p>Contributes to <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llam…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 12:55
<details open=""> <p>CUDA: size routed MoE MMQ N-tiles from typical expert width on RDNA3 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28552">#28552</a>)</p> <p>Recreated from <a class="issue-link js-issue-link" href="https://github.com/gg…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 11:47
<details open=""> <p>CUDA: replace GGML_FA_ALL_QUANTS with GGML_FA_QUANTS, more control over what is compiled (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28079">#28079</a>)</p> <ul> <li>CUDA: add configurable FA quant combinations</li> </…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 11:17
<details open=""> <p>args: officially deprecate --mmap|mlock|dio (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28334">#28334</a>)</p> <p>Signed-off-by: Aaron Teo <a href="mailto:[email protected] ">[email protected] </a></p> </details> <p><…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 09:56
<details open=""> <p>model: fix granite3 moe unknown parameter count (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28632">#28632</a>)</p> <p>Signed-off-by: Aaron Teo <a href="mailto:[email protected] ">[email protected] </a></p> </details> …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 09:24
<details open=""> <p>mtmd: propagate video ID to bitmap (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28601">#28601</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 07:47
<details open=""> <p>jinja: treat a null left operand of in as a plain lookup (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28620">#28620</a>)</p> <p>Templates that default an optional variable to none and then test its<br /> membership in …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 07:20
<details open=""> <p>vulkan: add dedicated iq4_xs mat-vec shader (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28426">#28426</a>)</p> <ul> <li>vulkan: add dedicated iq4_xs mat-vec shader</li> </ul> <p>Dedicated mul_mat_vec_iq4_xs for the dm…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 06:40
<details open=""> <p>vulkan: add f16 B-type matmul pipelines and warp tile size tuning for Intel coopmat1 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27471">#27471</a>)</p> <ul> <li> <p>vulkan: add f16 B-type matmul pipelines and warp til…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 05:27
<details open=""> <p>tests : use less threads for data initialization (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28325">#28325</a>)</p> <ul> <li> <p>tests : use 1 thread for data initialization</p> </li> <li> <p>cont : scale threads with…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-09 02:03
<details open=""> <p>Add IQ type handling for MoE (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28476">#28476</a>)</p> <p>Co-authored-by: cwriter <a href="mailto:cwriter@localhost">cwriter@localhost</a></p> </details> <p><strong>Website:</s…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 17:33
<details open=""> <p>llama: disable lazy tensor loading by default on iGPUs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28326">#28326</a>)</p> <ul> <li> <p>llama: add lazy mode auto, fix iGPU regression</p> </li> <li> <p>revert changes ex…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 15:14
<details open=""> <p>Revert "ggml-cuda : restore prop.integrated on HIP builds (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24233">#24233</a>)" (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28604">#2…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 14:33
<details open=""> <p>server : apply checkpoint min-step eviction only when the checkpoint list is full (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28302">#28302</a>)</p> <p>The spacing eviction in create_checkpoint() keeps the oldest chec…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 14:01
<details open=""> <p>metal : fix idle threads in mul_mv_iq3_xxs for ne00 < 1024 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28086">#28086</a>)</p> <ul> <li> <p>metal : fix half-idle simdgroup in kernel_mul_mv_iq3_xxs_f32 for ne00 < …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 12:36
<details open=""> <p>llama : add missing headers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28566">#28566</a>)</p> <ul> <li>fix compile-error: add missing header</li> </ul> <p>Bug: <a class="issue-link js-issue-link" href="https://github…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 12:08
<details open=""> <p>vulkan : fuse UNARY(GELU|SIGMOID|SILU|SOFTPLUS) + MUL (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27220">#27220</a>)</p> <ul> <li> <p>vulkan : fuse UNARY(SIGMOID|SILU|SOFTPLUS) + MUL</p> </li> <li> <p>vulkan : fuse UN…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 11:38
<details open=""> <p>Fix Vulkan-Hpp handle usage on 32-bit targets. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/22892">#22892</a>)</p> <p>On 32-bit platforms, Vulkan non-dispatchable handles such as VkBuffer are<br /> represented as uint6…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 10:50
<details open=""> <p>chat : split specialized parsers into common/parsers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27764">#27764</a>)</p> <ul> <li>chat : split specialized parsers into common/parsers</li> </ul> <p>Move the 14 dedicated…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 10:06
<details open=""> <p>opencl: properly handle non-contiguous inputs to conv2d (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28503">#28503</a>)</p> <ul> <li> <p>opencl: fix conv2d non-contiguous strides</p> </li> <li> <p>opencl: format</p> </…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 04:00
<details open=""> <p>model : support Kimi-K3 recurrent-state rollback (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28466">#28466</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-08 00:42
<details open=""> <p>hexagon: add RELU and LEAKY_RELU ops (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28585">#28585</a>)</p> <ul> <li> <p>hexagon: add RELU op</p> </li> <li> <p>hexagon: add LEAKY_RELU op too</p> </li> </ul> </details> <p>…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 20:05
<details open=""> <p>tests : initialize the L2_NORM batch array (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28553">#28553</a>)</p> <ul> <li>tests: bind the L2_NORM batch count to a local</li> </ul> <p>GCC cannot prove the loop fills norms…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 19:30
<details open=""> <p>vulkan: add DeepSeek-V4 hyper-connection fused ops (DSV4_HC_COMB/PRE/POST) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26578">#26578</a>)</p> <ul> <li>vulkan: add DeepSeek-V4 hyper-connection fused ops (DSV4_HC_COMB/P…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 18:16
<details open=""> <p>ggml: add gfx90c HIP support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26454">#26454</a>)</p> <ul> <li> <p>ggml: add gfx90c HIP support</p> </li> <li> <p>ggml: make gfx90c HIP support compliant with specifications</…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 14:33
<details open=""> <p>CUDA: branchless Q4_K/Q5_K unpack to speed up mmvq, L2 prefetch on DGX Spark (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26705">#26705</a>)</p> <ul> <li> <p>Update Q4_K and Q5_K to use branchless computation, which st…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 11:16
<details open=""> <p>vulkan: support type-aligned GET_ROWS (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28253">#28253</a>)</p> <ul> <li>vulkan: fall back to CPU for GET_ROWS with misaligned offsets</li> </ul> <p>The Vulkan GET_ROWS shader …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 08:37
<details open=""> <p>caps : recheck typed content if template checks for string (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28511">#28511</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofol…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 08:11
<details open=""> <p>ggml-cuda: fix divergent barrier in f16 flash attention (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27870">#27870</a>)</p> <ul> <li> <p>ggml-cuda: fix divergent barrier in f16 flash attention</p> </li> <li> <p>ggml-cu…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 07:20
<details open=""> <p>ggml: allow backend inputs to not create another split (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28387">#28387</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow"…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 06:51
<details open=""> <p>vulkan: rms_norm fusion opportunities (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28024">#28024</a>)</p> <p>Support RMS_NORM + MUL + ADD (+ MUL) and RMS_NORM + VIEW + SET_ROWS.<br /> Extend ROPE + VIEW + SET_ROWS to s…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-07 05:21
<details open=""> <p>vulkan: add TQ1_0 support (mm, mat-vec, mat-vec-id, dequant, get_rows) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27765">#27765</a>)</p> <ul> <li> <p>vulkan: add TQ1_0 support (mm, mat-vec, dequant, get_rows)</p> </l…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-06 23:30
<details open=""> <p>convert : add <code>--fuse-qkv</code> flag to fuse Q/K/V into QKV during HF-to-GGUF conversion (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/22780">#22780</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-06 23:06
<details open=""> <p>models : fix GDN normalization from <code>max</code> to <code>rsqrt</code> (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28068">#28068</a>)</p> <ul> <li>models: use flash-linear-attention's l2norm for gated delta net q/…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-06 22:11
<details open=""> <p>[Model] Support for Spark2_5ForCausalLM implementation (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27868">#27868</a>)</p> <ul> <li>Add Spark3 Model</li> <li>rename spark3 -> spark2_5</li> </ul> <p>Co-authored-by: S…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-06 20:52
<details open=""> <p>opencl: properly choose weights pack for q4_K, q5_K mul_mat (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28402">#28402</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofo…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-06 16:42
<details open=""> <p>cuda: fixes races in mmid and mmf (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28475">#28475</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-06 13:24
<details open=""> <p>grammar : fix max repetition threshold (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28469">#28469</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-06 11:59
<details open=""> <p>common: add --log-jsonl (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28437">#28437</a>)</p> <ul> <li> <p>common: add --log-jsonl</p> </li> <li> <p>rename unknown to none</p> </li> </ul> </details> <p><strong>Website:</…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-06 11:32
<details open=""> <p>ui : embed assets directly with CMake (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28445">#28445</a>)</p> <p>Remove the build-time C++ helper and external gzip dependency,<br /> simplifying cross-compilation. Keep the …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-06 10:01
<details open=""> <p>metal : add remaining fa-vec tunings for M2 Max (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28458">#28458</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:…
llama.cpp — Releases
TIER_1
(SO)
·
JohannesGaessler
·
2026-09-05 20:42
<p>Github: limit blank issues to maintainers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28435">#28435</a>)</p>
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-05 10:38
<details open=""> <p>metal : fix memory leak in early return (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28399">#28399</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-05 09:28
<details open=""> <p>sycl: attribute device allocations by site (GGML_SYCL_MEMTRACE) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27631">#27631</a>)</p> <p>define two new environment variables to better understand how much<br /> memory is …
llama.cpp — Releases
TIER_1
(SO)
·
philip-jingxin
·
2026-09-05 02:37
<p>sycl : fix test-backend-ops CI break && restore Kronecker product FWH…</p>
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-04 20:00
<details open=""> <p>metal : add remaining fa-vec tunings for M3 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28396">#28396</a>)</p> <ul> <li> <p>addition of m3 in fa_vec_tuned_table</p> </li> <li> <p>adding q4_0,q4_1,q5_0,q5_1 in ggml-met…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-04 19:32
<details open=""> <p>opencl: extend the elementwise and data‐movement op coverage (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27633">#27633</a>)</p> <ul> <li>opencl: add extended elementwise unary ops (sgn, step, elu, hardswish, hardsigmo…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-04 19:05
<details open=""> <p>opencl: add Adreno xmem SDPA path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26331">#26331</a>)</p> <ul> <li>opencl: add Adreno xmem SDPA path</li> </ul> <p>Assisted-by: Codex</p> <ul> <li> <p>Removed the Adreno-spec…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-09-04 17:57
<details open=""> <p>llama.cpp : bump version to 0.4.0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28386">#28386</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…