llama.cpp project releases pre-release updates with various fixes and optimizations
ByPulseAugur Editorial·[12 sources]·
The llama.cpp project has released several pre-release updates, including b11514 which addresses Musa FWHT fixes and platform attestations for various operating systems. Other releases like b11512 and b11511 focus on model fixes and CUDA optimizations, respectively. Notably, b11509 includes a fix for CUDA CCCL version guards, and b11505 resolves issues with Vulkan TOP_K for infinite and NaN inputs. The updates also feature contributions from various developers and integrations with libraries like cpp-httplib.
AI
IMPACT
These updates improve the performance and stability of running LLMs on various platforms.
RANK_REASON
The cluster consists of multiple pre-release updates to the llama.cpp project, which is a software tool for running large language models.
<details open=""> <p>CUDA : looped PAD kernel for more than 65535 rows or slices (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30147">#30147</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofo…
<details open=""> <p>CUDA: fix CCCL version guard breaking on major version rollover (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29453">#29453</a>)</p> <ul> <li>CUDA: fix CCCL version guard breaking on major version rollover</li> </ul> <p…
<details open=""> <p>vulkan : fix TOP_K for +inf/NaN inputs and k = 1 on negative values (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30107">#30107</a>)</p> <p>The bucket search in topk_nary_search.comp started from the range<br /> [0, 0xF…
<p>vulkan: extend sparse FA support to coopmat2 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/30003">#30003</a>)</p>