llama.cpp releases bring ZDNN backend, Vulkan tuning, and Hexagon support
ByPulseAugur Editorial·[47 sources]·
The llama.cpp project has released several updates, including version b11269 which adds a ZDNN backend build and various CI improvements. Other recent releases focus on specific optimizations and bug fixes, such as improving OpenCL kernels for Adreno GPUs, tuning Vulkan shaders for Intel hardware, and enhancing FP32 dot product accumulation for AVX512-FP16. These updates also include adjustments to vocabulary handling for PLaMo models and support for FP32 GELU_ERF and GEGLU_ERF on Hexagon processors.
AI
IMPACT
Ongoing performance optimizations and backend support for llama.cpp, enabling broader hardware compatibility and efficiency.
RANK_REASON
This is a series of software releases for an open-source project, not a frontier model release or significant industry event.
<details open=""> <p>cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29722">#29722</a>)</p> <ul> <li>cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast</li> </ul…
<details open=""> <p>ci : fix Models Backend Check by shortening the hrm_text fixture (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29744">#29744</a>)</p> <p>The fixture recycles its two blocks over 8 cache slots, so the fp16<br /> error bu…
<details open=""> <p>cpu: accept BF16 in src1 of mul_mat (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28937">#28937</a>)</p> <ul> <li>cpu: accept BF16 in src1 of mul_mat</li> </ul> <p>ggml_conv_1d_dw builds its im2col in F32 when the kerne…
<details open=""> <p>openvino: serve GET_ROWS on a weight view from the base Constant (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28381">#28381</a>)</p> <ul> <li>openvino: serve GET_ROWS on a weight view from the base Constant</li> </ul> …
<details open=""> <p>musa : define <strong>CUDA_ARCH</strong> for device passes (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29508">#29508</a>)</p> <p>The MUSA vendor header never defined <strong>CUDA_ARCH</strong>, so every architecture<b…
<details open=""> <p>hexagon: optimize concat op (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29673">#29673</a>)</p> <ul> <li>hex-concat: reduce pkts in gather/transpose hot loop</li> </ul> <p>gather directly into dst buffer, use special i…
<details open=""> <p>vulkan : Load F32 A matrix 2 at a time when its 2-aligned (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29254">#29254</a>)</p> <p>It turns out Intel doesn't particularly like loading F32s one at a<br /> time and we alre…
<details open=""> <p>vulkan: MOE aware mat_mul_id tile selection (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29182">#29182</a>)</p> <p>mut_mul_id selected its matmul tile with total token count.<br /> For MoE dispatch grid the true N per …
<details open=""> <p>ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29504">#29504</a>)</p> <ul> <li> <p>fix c++ odr by properly using GGML_COMMON_DECL_CPP</p> </li> <li> <p>using actu…
<details open=""> <p>vocab : keep NORMAL in PLaMo-2 and PLaMo-3 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29580">#29580</a>)</p> <ul> <li>vocab : keep NORMAL in PLaMo-2 and PLaMo-3</li> </ul> <p>The PLaMo-2 and PLaMo-3 vocabularies mark…
<details open=""> <p>server : remove the built-in UI's service worker when the UI is not served (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29565">#29565</a>)</p> <p>With --path or --no-ui, /sw.js returned 404, and a 404 does not remove a…
<details open=""> <p>ggml : collect all input tensors into graph_inputs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29634">#29634</a>)</p> <p>graph_inputs was populated while splitting the graph, so it only<br /> contained the inputs that…
<details open=""> <p>common : use fs::path for cache dirs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29595">#29595</a>)</p> <ul> <li>Avoid useless string conversions on Windows.</li> <li>No need for BSD or emscripten special cases.</li> …
<details open=""> <p>models: pad on the left with ggml_pad_ext (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29567">#29567</a>)</p> <ul> <li>models: pad on the left with ggml_pad_ext</li> </ul> <p>The Parakeet, LFM2-Audio, Granite Speech an…