llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-10 10:53
<details open=""> <p>model-saver : fix expert shared/chunk FFN length key clobber (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26693">#26693</a>)</p> <p>The saver called add_kv with LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH twice, the<br />…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-10 08:19
<details open=""> <p>ggml-webgpu : refactor several wgsl files and simplify flash_attn wgsl. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26134">#26134</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.ap…
llama.cpp — Releases
TIER_1
(SO)
·
ServeurpersoCom
·
2026-08-09 19:20
<p>ui: degrade the working directory picker when file search is off (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26">#26</a>…</p>
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-09 11:23
<details open=""> <p>ggml-cpu : fix missing Q5_0 dispatch in SpaceMiT backend (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26792">#26792</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollo…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-09 10:50
<details open=""> <p>ci: rm <code>GGML_HIP_ROCWMMA_FATTN</code> (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26760">#26760</a>)</p> <p>Signed-off-by: Aaron Teo <a href="mailto:[email protected] ">[email protected] </a></p> </details> <p><s…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-08 23:28
<details open=""> <p>server: report the isolate working directory from get_info (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26773">#26773</a>)</p> <ul> <li>server: report the isolate working directory from get_info</li> </ul> <p>Without a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-08 17:22
<details open=""> <p>CUDA: fuse rms_norm + mul + rope (+ view + set_rows) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26767">#26767</a>)</p> <ul> <li> <p>CUDA: fuse rms_norm + mul + rope (+ view + set_rows)</p> </li> <li> <p>tests: add br…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-07 14:07
<details open=""> <p>Mitigate crashing issue on Windows MSYS2 UCRT64 environment (GCC 16.1.0) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26555">#26555</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.a…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-07 13:22
<details open=""> <p>sycl: fix UE4M3 parsing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25608">#25608</a>)</p> <p>The NVFP4 quantization format stores a scaling factor for every group of<br /> 16 weights, packed into a single UE4M3 byte.…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-07 12:35
<details open=""> <p>sycl: *glu flat path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26354">#26354</a>)</p> <ul> <li>tests: add SWIGLU perf cases</li> </ul> <p>perf mode had no GLU coverage. Adds SWIGLU at 17408 columns, 512 and<br /> 20…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-07 12:05
<details open=""> <p>sycl : Support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC_POST,DSV4_HC_PRE (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26568">#26568</a>)</p> <ul> <li> <p>support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-07 11:17
<details open=""> <p>sycl : fix error Error OP FLASH_ATTN_EXT on arc770 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26441">#26441</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">htt…
llama.cpp — Releases
TIER_1
(SO)
·
github-actions[bot]
·
2026-08-07 10:03
<details open=""> <p>cuda: fix warnings for unused variable/function (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26688">#26688</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:…