llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 22:06
<details open=""> <p>vulkan: add POOL_1D op (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25431">#25431</a>)</p> <ul> <li>vulkan : add pool1d push constants and pipeline field</li> </ul> <p>Declared data structures needed for POOL1D OP, whi…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 21:33
<details open=""> <p>vulkan: Introduce driver version check for Windows Intel GPU to mitigate crashing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25192">#25192</a>)</p> <ul> <li>Removed crash guard for Intel</li> </ul> <p>Crash fixed fro…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 20:47
<details open=""> <p>mtmd: add n_embd_head (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26342">#26342</a>)</p> <p>Co-authored-by: Daniel Han <a href="mailto:[email protected] ">[email protected] </a></p> </details> <p><strong>Website:</s…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 20:12
<details open=""> <p>Support rotated kv cache quant (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26180">#26180</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 19:27
<details open=""> <p>llama : load MTP tensors only if they are really used (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26296">#26296</a>)</p> <ul> <li> <p>llama : load MTP tensors only if they are really used</p> </li> <li> <p>llama : ski…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 18:53
<details open=""> <p>vulkan: update vulkan sdk to 1.4.357.0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26303">#26303</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 18:11
<details open=""> <p>server: correct accepted tokens when need draft token replay (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26320">#26320</a>)</p> <ul> <li> <p>spec: correct accepted tokens when need draft token replay</p> </li> <li> <p…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 17:30
<details open=""> <p>cuda: extract Q2_0 elements via __byte_perm (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25603">#25603</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://ll…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 16:05
<details open=""> <p>SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt proc… (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25025">#25025</a>)</p> <ul> <li> <p>SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt p…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 15:14
<details open=""> <p>[SYCL] support the missed types in cpy (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26005">#26005</a>)</p> <ul> <li> <p>support the missed types in cpy</p> </li> <li> <p>use correct funct</p> </li> <li> <p>rm unused co…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 13:28
<details open=""> <p>ggml-zendnn : group matmul direct API for mul_mat_id (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25918">#25918</a>)</p> <ul> <li> <p>ggml-zendnn : group matmul API for mul_mat_id</p> </li> <li> <p>ggml-zendnn : scale …
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 12:46
<details open=""> <p>sycl : support dev2dev memcpy by DEV2DEV_MEMCPY_FORWARD (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26234">#26234</a>)</p> <p>Co-authored-by: Neo Zhang Jianyu <a href="mailto:[email protected] ">jianyu.zhang@intel…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 12:16
<details open=""> <p>[SYCL] Support q2 mul_mat (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26231">#26231</a>)</p> <ul> <li> <p>support q2_0 in mul_mat</p> </li> <li> <p>support more q2_0 case</p> </li> </ul> </details> <p><strong>Website:…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 11:43
<details open=""> <p>sycl: fuse RMS_NORM + MUL (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26015">#26015</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> <…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-07-31 09:47
<details open=""> <p>ggml-webgpu: improve flash_attn_vec for quantized KV at long contexts (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25956">#25956</a>)</p> <ul> <li> <p>improve fa of quantized kv cache</p> </li> <li> <p>Fix some bugs an…