llama.cpp releases multiple updates with cross-platform optimizations
ByPulseAugur Editorial·[76 sources]·
The llama.cpp project has released several updates, including versions b10106, b10105, b10108, b10099, b10098, b10094, b10093, b10092, b10091, and b10103. These releases introduce various improvements and fixes across different platforms and hardware accelerators. Notable updates include enhancements for CUDA, Metal, Hexagon, and OpenVINO, as well as fixes for specific model templates like DeepSeekv4 and improvements to argument parsing and memory management.
AI
IMPACT
Ongoing improvements to a popular open-source inference engine enhance its performance and compatibility across diverse hardware and operating systems.
RANK_REASON
The cluster consists of multiple release notes for the llama.cpp project, detailing software updates and bug fixes.
<details open=""> <p>CUDA: Improve NVFP4 W4A4 activation quantization (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25730">#25730</a>)</p> <ul> <li>Squash history before conflict-resolution during rebase on master</li> </ul> <p>WIP commit</…
<details open=""> <p>common: infer the speculative type from the draft repo sidecars (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25989">#25989</a>)</p> <p>With -hfd pointing to a repo that ships mtp-/dflash-/eagle3- sidecars<br /> and no …
<p>metal : add f16 type support to leaky relu (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25981">#25981</a>)</p>
<details open=""> <p>mtmd : use align_corners for qwen3vl vision position embedding interpolation (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25781">#25781</a>)</p> <p>The Qwen3-VL learned position embedding is interpolated to the runtime…
<details open=""> <p>kleidiai : warn once when a weight type has no KleidiAI kernel (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25701">#25701</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="n…
<details open=""> <p>common: resolve draft repo to its requested sidecar (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25955">#25955</a>)</p> <p>With -hfd pointing to a repo shipping speculative sidecars, the draft<br /> resolved to the mai…
<details open=""> <p>server: return 400 instead of 500 on validation error with X-Conversation-Id (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25760">#25760</a>)</p> <ul> <li>server: return 400 instead of 500 on validation error with X-Con…
<p>llama-arch: fix DeepSeek4 APE tensor op (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25945">#25945</a>)</p>
<details open=""> <p>vulkan: Refactor vk_queue to use per-instance mutexes and unique handles (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23570">#23570</a>)</p> <ul> <li> <p>Refactor vk_queue to use per-instance mutexes and unique handles…
<details open=""> <p>opencl: Support broadcast for Adreno MUL_MAT and honor <code>view_offs</code> for Adreno Q8_0 MUL_MAT for llama-server multi-stream (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25910">#25910</a>)</p> <ul> <li> <p>openc…
<details open=""> <p>opencl: load and use <code>kernel_gemm_moe_q6_k_f32_ns</code> from bin kernel lib (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25797">#25797</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https:…
<details open=""> <p>tests : initialize all tensors in test_dsv4_hc to avoid NaNs in sentinel tensors (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25822">#25822</a>)</p> <p>Co-authored-by: Stanisław Szymczyk <a href="mailto:sszymczy@gmail.…
<details open=""> <p>vulkan: Support Q2_0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25430">#25430</a>)</p> <ul> <li>vulkan: Support Q2_0</li> </ul> <p>The backend perf tests for mat-vec-mul weren't very good at first (worse than<br /> q…
<details open=""> <p>hexagon: L2 cache handling rework (dirty bit tracking with lazy flushing) and more MUL_MAT updates (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25762">#25762</a>)</p> <ul> <li> <p>hex-mm: fix artificial limit in the so…
<details open=""> <p>kleidiai: Add SME vs SME2 distinction in kernel dispatch (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25478">#25478</a>)</p> <p>The current integration treats SME as a single capability (CPU_FEATURE_SME)<br /> with no …
<details open=""> <p>vulkan: when using transfer queue for async copies, sync on event_wait to avoid race (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25229">#25229</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="htt…
<details open=""> <p>opencl: disable FA and MoE weights repack to work around compiler issues for Adreno 850 GPU (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25745">#25745</a>)</p> <ul> <li> <p>opencl: workaround for A850 compiler compat</…
<details open=""> <p>sycl: Increase minimum buffer size for USM system allocations (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25525">#25525</a>)</p> <p>Raise the threshold for minimum buffer size from 1 GiB to 4 GiB, based<br /> on real-…
<details open=""> <p>opencl: do not use <code>clCreateBufferWithProperties</code> when targeting CL 2.x (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25673">#25673</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="htt…
<details open=""> <p>opencl: fix a dp4a bug for devices where cl_khr_integer_dot_product is unavailable (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25639">#25639</a>)</p> <ul> <li> <p>opencl: do not fail backend init on devices without cl…
<details open=""> <p>vulkan/cpu: Support f16 as SET_ROWS src. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25432">#25432</a>)</p> <ul> <li>vulkan/cpu: Support f16 as SET_ROWS src.</li> </ul> <p>This adds full support for f16 SET_ROWS (equi…
<details open=""> <p>ggml : add a set of functions for checking contiguity of inner tensor dimensions (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25650">#25650</a>)</p> <p>Co-authored-by: Stanisław Szymczyk <a href="mailto:sszymczy@gmail.…
<details open=""> <p>tests: export-graph-ops: exit gracefully when called w/o arguments (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25619">#25619</a>)</p> <p>Fixes a segfault when <code>test-export-graph-ops</code> is called without any<b…