llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-27 00:22
<details open=""> <p>mtmd: Add Vision Support for Minimax-M3 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25113">#25113</a>)</p> <ul> <li>Add preliminary MiniMax-M3 support</li> </ul> <p>Text-only port that re-uses existing components: Min…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-26 23:05
<details open=""> <p>mtmd: fix android build (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26150">#26150</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> </u…
llama.cpp — Releases
TIER_1
(SO)
·
ServeurpersoCom
·
2026-07-26 04:51
<p>ui: fix context gauge card regressions and land at the conversation e…</p>
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-24 14:02
<details open=""> <p>hexagon: fix Windows crash when op_poll is enabled (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26029">#26029</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">htt…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-24 11:36
<details open=""> <p>CUDA: fix external compilation of q1_0 MMQ (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25778">#25778</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://lla…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-24 09:35
<details open=""> <p>args: refactor mlock/mmap/directio into load-mode (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/20834">#20834</a>)</p> <ul> <li>args: overhaul mmap/mlock/dio into single arg</li> </ul> <p>Signed-off-by: Aaron Teo <a hre…
llama.cpp — Releases
TIER_1
(SO)
·
max-krasnyansky
·
2026-07-24 02:13
<p>hexagon: further improved pipeline of the core bits (L2, DMA, MM, FA)…</p>
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-23 23:35
<details open=""> <p>CUDA: Improve NVFP4 W4A4 activation quantization (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25730">#25730</a>)</p> <ul> <li>Squash history before conflict-resolution during rebase on master</li> </ul> <p>WIP commit</…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-23 20:13
<details open=""> <p>hexagon: activation ops update (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25974">#25974</a>)</p> <ul> <li> <p>hex-geglu: optimized all-in-one geglu microkernel</p> </li> <li> <p>hex-geglu: enable non-contiguous src a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-23 18:29
<details open=""> <p>common: infer the speculative type from the draft repo sidecars (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25989">#25989</a>)</p> <p>With -hfd pointing to a repo that ships mtp-/dflash-/eagle3- sidecars<br /> and no …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-23 17:21
<details open=""> <p>Fix DeepSeek4 crafted template (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25414">#25414</a>)</p> <ul> <li> <p>chat: fix DS4 template to explicitly follow reference behavior</p> </li> <li> <p>Support DeepSeekv4 flag (…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-23 05:32
<details open=""> <p>ggml: enable PowerPC backend variants on AIX (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25983">#25983</a>)</p> <ul> <li>ggml: enable PowerPC backend variants on AIX</li> </ul> <p>Allow the PowerPC CPU backend variant…
llama.cpp — Releases
TIER_1
(SO)
·
iliailmer
·
2026-07-23 03:45
<p>metal : add f16 type support to leaky relu (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25981">#25981</a>)</p>
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 21:13
<details open=""> <p>ci : fix SYCL package shared library lookup (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25987">#25987</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://ll…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 20:25
<details open=""> <p>webgpu : add CONV_2D_DW (depthwise conv2d) kernel (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25847">#25847</a>)</p> <ul> <li>webgpu : add CONV_2D_DW (depthwise conv2d) kernel</li> </ul> <p>Implement GGML_OP_CONV_2D_D…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 19:47
<details open=""> <p>cuda: GET_ROWS quants (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25962">#25962</a>)</p> <ul> <li>cuda: add k-quant support to GET_ROWS</li> </ul> <p>Device-side embedding lookups require GET_ROWS to handle the k-quan…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 11:05
<details open=""> <p>Add support for Laguna XS.2 & M.1 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25165">#25165</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.ap…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 09:54
<details open=""> <p>mtmd : use align_corners for qwen3vl vision position embedding interpolation (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25781">#25781</a>)</p> <p>The Qwen3-VL learned position embedding is interpolated to the runtime…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 09:14
<details open=""> <p>hexagon: check tensor type when reusing descriptors (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25968">#25968</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">ht…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 08:35
<details open=""> <p>cuda: add sqrt_softplus in topk-moe for dsv4 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25896">#25896</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://l…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 06:29
<details open=""> <p>kleidiai : warn once when a weight type has no KleidiAI kernel (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25701">#25701</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="n…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 05:44
<details open=""> <p>common: resolve draft repo to its requested sidecar (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25955">#25955</a>)</p> <p>With -hfd pointing to a repo shipping speculative sidecars, the draft<br /> resolved to the mai…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 05:07
<details open=""> <p>server: return 400 instead of 500 on validation error with X-Conversation-Id (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25760">#25760</a>)</p> <ul> <li>server: return 400 instead of 500 on validation error with X-Con…
llama.cpp — Releases
TIER_1
(SO)
·
helanfxz
·
2026-07-22 02:55
<p>llama-arch: fix DeepSeek4 APE tensor op (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25945">#25945</a>)</p>
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-22 02:05
<details open=""> <p>server : properly handle null llama_context (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25868">#25868</a>)</p> <p>Co-authored-by: Stanisław Szymczyk <a href="mailto:[email protected] ">[email protected] </a></p> </det…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-21 22:40
<details open=""> <p>vulkan: Refactor vk_queue to use per-instance mutexes and unique handles (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23570">#23570</a>)</p> <ul> <li> <p>Refactor vk_queue to use per-instance mutexes and unique handles…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-21 21:55
<details open=""> <p>ggml-openvino: Add GGML_BACKEND_DL_IMPL invocation for OpenVINO backend (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25795">#25795</a>)</p> <p>This adds the missing <code>GGML_BACKEND_DL_IMPL()</code> macro invocation,…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-21 15:54
<details open=""> <p>CUDA: vectorize same-type get_rows with int4 copy (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25929">#25929</a>)</p> <p>k_get_rows_float did a scalar one-element-per-thread copy and recomputed the<br /> row-invariant …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-20 06:31
<details open=""> <p>opencl: Support broadcast for Adreno MUL_MAT and honor <code>view_offs</code> for Adreno Q8_0 MUL_MAT for llama-server multi-stream (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25910">#25910</a>)</p> <ul> <li> <p>openc…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-18 13:38
<details open=""> <p>model: rotate injected K/V cache for DFlash (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25823">#25823</a>)</p> <ul> <li> <p>dflash: rotate injected K/V cache when using K/V quantization</p> </li> <li> <p>Update src/mo…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-18 12:14
<details open=""> <p>llama-quant : exclude i32 ffn_gate_tid2eid routing table from quantization (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25787">#25787</a>)</p> <p>DeepSeek-V4's ffn_gate_tid2eid tensor is an i32 token-id -> expert-id…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-17 23:13
<details open=""> <p>opencl: load and use <code>kernel_gemm_moe_q6_k_f32_ns</code> from bin kernel lib (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25797">#25797</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https:…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-17 15:42
<details open=""> <p>opencl: transpose q4_K noshuffle scales for coalesced reads (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25805">#25805</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofo…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-17 15:00
<details open=""> <p>sync : ggml</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> </ul> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download/b1006…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-17 14:20
<details open=""> <p>tests : initialize all tensors in test_dsv4_hc to avoid NaNs in sentinel tensors (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25822">#25822</a>)</p> <p>Co-authored-by: Stanisław Szymczyk <a href="mailto:sszymczy@gmail.…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-17 12:24
<details open=""> <p>ggml-blas: default hadamard mul_mat to cpu routine (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25710">#25710</a>)</p> <p>Signed-off-by: Aaron Teo <a href="mailto:[email protected] ">[email protected] </a></p> </detail…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-17 11:45
<details open=""> <p>vulkan: Support Q2_0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25430">#25430</a>)</p> <ul> <li>vulkan: Support Q2_0</li> </ul> <p>The backend perf tests for mat-vec-mul weren't very good at first (worse than<br /> q…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-17 10:24
<details open=""> <p>sycl: fix row calculation when K_QUANTS_PER_ITERATION is 1 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25690">#25690</a>)</p> <ul> <li>sycl: fix incorrect row calculation when K_QUANTS_PER_ITERATION=1</li> </ul> <p>Si…
llama.cpp — Releases
TIER_1
(SO)
·
Gezahegne
·
2026-07-17 05:13
<p>opencl: add ABS op (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25115">#25115</a>)</p>
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-17 00:30
<details open=""> <p>docs: added a note about using OpenCl with Adreno 810 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25786">#25786</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 21:25
<details open=""> <p>hexagon: L2 cache handling rework (dirty bit tracking with lazy flushing) and more MUL_MAT updates (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25762">#25762</a>)</p> <ul> <li> <p>hex-mm: fix artificial limit in the so…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 20:51
<details open=""> <p>kleidiai: Add SME vs SME2 distinction in kernel dispatch (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25478">#25478</a>)</p> <p>The current integration treats SME as a single capability (CPU_FEATURE_SME)<br /> with no …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 20:10
<details open=""> <p>vulkan: when using transfer queue for async copies, sync on event_wait to avoid race (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25229">#25229</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="htt…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 19:24
<details open=""> <p>TP: fix Phi3, Bert, Plamo2/3, ChatGLM (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25536">#25536</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.ap…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 18:46
<details open=""> <p>vendor: update BoringSSL to 0.20260713.0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25624">#25624</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 17:08
<details open=""> <p>tests: actually exercise <code>test-recurrent-state-rollback</code> (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25758">#25758</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" r…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 16:11
<details open=""> <p>server : allow text-only slot save/restore with mtmd (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25076">#25076</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">h…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 14:04
<details open=""> <p>CUDA: Support CUDA Virtual Devices (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25228">#25228</a>)</p> <ul> <li> <p>support cuda virtual devices</p> </li> <li> <p>disable NCCL path when virtual devices are used</p> </l…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 11:24
<details open=""> <p>Enable CUDA graphs on volta+turing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25749">#25749</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 10:38
<details open=""> <p>server: Ignore empty / non-existing <code>Origin</code> headers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25756">#25756</a>)</p> <p>Otherwise this gives lots of unnecessary warnings:</p> <p>W srv operator(): (CORS) …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 09:57
<details open=""> <p>ggml-cuda : restore prop.integrated on HIP builds (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24233">#24233</a>)</p> <p>PR <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/16308">#1…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 06:59
<details open=""> <p>ci : add official website link to release notes (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25728">#25728</a>)</p> <p>Assisted-by: pi:llama.cpp/Qwen3.6-27B</p> </details> <p><strong>Website:</strong></p> <ul> <li><a h…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 06:17
<details open=""> <p>quant : allow using manual tensor types with --pure (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25716">#25716</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 05:32
<details open=""> <p>opencl: disable FA and MoE weights repack to work around compiler issues for Adreno 850 GPU (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25745">#25745</a>)</p> <ul> <li> <p>opencl: workaround for A850 compiler compat</…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-16 05:00
<details open=""> <p>cuda: extract Q1_0 elements via __byte_perm (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25628">#25628</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/rele…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 20:37
<details open=""> <p>opencl: exclude some moe kernels on Adreno a7x (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25698">#25698</a>)</p> <ul> <li>opencl: exclude Adreno A7x from using Adreno MoE kernels</li> </ul> <p>Some compilers for A7x …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 19:53
<details open=""> <p>cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25545">#25545</a>)</p> <ul> <li> <p>cuda : CUDA GGML_OP_LIGHTNING_INDEXER implemen…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 19:17
<details open=""> <p>tokenize : drop --stdin mutual-exclusion check (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25672">#25672</a>)</p> <p>match cli and completion, which don't enforce it</p> </details> <p><strong>macOS/iOS:</strong></p> <…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 16:12
<details open=""> <p>cuda : relax tensor contiguity requirements for quantized concat (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25678">#25678</a>)</p> <ul> <li> <p>cuda : relax tensor contiguity requirements for quantized concat</p> </l…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 14:22
<details open=""> <p>DeepseekV4: reduce graph splits (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25702">#25702</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/downloa…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 13:39
<details open=""> <p>sycl : fix get_rows Q2_K, Q4_K, Q5_K (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25656">#25656</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/do…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 13:03
<details open=""> <p>sycl : support kernel type fp16 for conv2d_dw (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25653">#25653</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/re…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 12:23
<details open=""> <p>sycl : implement xielu op (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25550">#25550</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download/b100…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 11:39
<details open=""> <p>sycl: Increase minimum buffer size for USM system allocations (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25525">#25525</a>)</p> <p>Raise the threshold for minimum buffer size from 1 GiB to 4 GiB, based<br /> on real-…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 09:14
<details open=""> <p>[SYCL] Flash Attention with XMX engine via oneDNN (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25222">#25222</a>)</p> <ul> <li> <p>[SYCL] F16 (default) Flash Attention with XMX engine via oneDNN graph API; Qwen3.6-27b-…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-15 03:26
<details open=""> <p>opencl: do not use <code>clCreateBufferWithProperties</code> when targeting CL 2.x (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25673">#25673</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="htt…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 21:36
<details open=""> <p>hexagon: fix hmx-queue signal enum-narrowing problem (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25677">#25677</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 20:57
<details open=""> <p>server : refactor prompt cache state ownership (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25649">#25649</a>)</p> <ul> <li> <p>server : clear checkpoints upon prompt clear</p> </li> <li> <p>server : move the prompt st…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 20:27
<details open=""> <p>server: add --cors-* options (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25655">#25655</a>)</p> <ul> <li> <p>server: add --cors-* options</p> </li> <li> <p>add special "localhost" value</p> </li> <li> <p>add tests</p>…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 19:44
<details open=""> <p>opencl: fix a dp4a bug for devices where cl_khr_integer_dot_product is unavailable (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25639">#25639</a>)</p> <ul> <li> <p>opencl: do not fail backend init on devices without cl…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 18:59
<details open=""> <p>DeepseekV4: fix seq_rm (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25588">#25588</a>)</p> <ul> <li> <p>DeepseekV4: fix seq_rm</p> </li> <li> <p>implement proper seq_cp</p> </li> <li> <p>create actual update context</p…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 18:18
<details open=""> <p>vulkan/cpu: Support f16 as SET_ROWS src. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25432">#25432</a>)</p> <ul> <li>vulkan/cpu: Support f16 as SET_ROWS src.</li> </ul> <p>This adds full support for f16 SET_ROWS (equi…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 17:44
<details open=""> <p>tokenize : align usage by using common args (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25516">#25516</a>)</p> <p>Migrate the tokenize tool to common_params_parse, replacing its<br /> hand-rolled argv parsing, Windows…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 15:21
<details open=""> <p>ggml : add a set of functions for checking contiguity of inner tensor dimensions (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25650">#25650</a>)</p> <p>Co-authored-by: Stanisław Szymczyk <a href="mailto:sszymczy@gmail.…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 12:47
<details open=""> <p>tests: export-graph-ops: exit gracefully when called w/o arguments (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25619">#25619</a>)</p> <p>Fixes a segfault when <code>test-export-graph-ops</code> is called without any<b…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-14 12:07
<details open=""> <p>ggml: uniformize im2col dst_type for all conv ops (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23660">#23660</a>)</p> <ul> <li> <p>ggml: uniformize im2col dst_type for all conv ops</p> </li> <li> <p>Update ggml/src/ggm…