PulseAugur
中
实时 13:48:22

llama.cpp 发布引入 ZDNN 后端、Vulkan 调优和 Hexagon 支持

llama.cpp 项目发布了多项更新,包括添加 ZDNN 后端构建和各种 CI 改进的版本 b11269。其他近期版本侧重于特定优化和错误修复,例如改进 Adreno GPU 的 OpenCL 内核、为 Intel 硬件调优 Vulkan 着色器,以及增强 AVX512-FP16 的 FP32 点积累加。这些更新还包括对 PLaMo 模型词汇处理的调整,以及对 Hexagon 处理器上 FP32 GELU_ERF 和 GEGLU_ERF 的支持。 AI

影响 llama.cpp 的持续性能优化和后端支持,实现了更广泛的硬件兼容性和效率。

排序理由 这是一系列开源项目的软件发布,而非前沿模型发布或重大行业事件。

在 llama.cpp — Releases 阅读 →

AI 生成摘要 · Google Gemini · 来自 47 个来源。 我们如何撰写摘要 →

llama.cpp 发布引入 ZDNN 后端、Vulkan 调优和 Hexagon 支持

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是一系列开源项目的软件发布,而非前沿模型发布或重大行业事件。
Source corroboration
47 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
3 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+18 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [47]

  1. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11299

    <details open=""> <p>cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29722">#29722</a>)</p> <ul> <li>cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast</li> </ul…

  2. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11298

    <details open=""> <p>mimo : support dflash (convert + feature extraction) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29650">#29650</a>)</p> <ul> <li> <p>convert : update to support dflash</p> </li> <li> <p>cont : fix</p> </li> </ul> <p>C…

  3. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11297

    <details open=""> <p>jinja : support coerced array attributes (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29574">#29574</a>)</p> <ul> <li> <p>support coerced array attributes</p> </li> <li> <p>add tests</p> </li> </ul> </details> <p><stro…

  4. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11295

    <details open=""> <p>ci : fix Models Backend Check by shortening the hrm_text fixture (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29744">#29744</a>)</p> <p>The fixture recycles its two blocks over 8 cache slots, so the fp16<br /> error bu…

  5. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11294

    <details open=""> <p>llama: llama_prefetch_rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29599">#29599</a>)</p> <ul> <li> <p>llama: llama_prefetch_rows</p> </li> <li> <p>llama: support row prefetch on Windows</p> </li> </ul> <p>Apply t…

  6. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11293

    <details open=""> <p>ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29675">#29675</a>)</p> <ul> <li> <p>ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)</p> </li> <li> …

  7. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11292

    <details open=""> <p>cpu: accept BF16 in src1 of mul_mat (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28937">#28937</a>)</p> <ul> <li>cpu: accept BF16 in src1 of mul_mat</li> </ul> <p>ggml_conv_1d_dw builds its im2col in F32 when the kerne…

  8. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11284

    <details open=""> <p>openvino: serve GET_ROWS on a weight view from the base Constant (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28381">#28381</a>)</p> <ul> <li>openvino: serve GET_ROWS on a weight view from the base Constant</li> </ul> …

  9. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11282

    <details open=""> <p>musa : define <strong>CUDA_ARCH</strong> for device passes (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29508">#29508</a>)</p> <p>The MUSA vendor header never defined <strong>CUDA_ARCH</strong>, so every architecture<b…

  10. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11280

    <details open=""> <p>SYCL: reduce tensor allreduce sync with pinned host buffers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29604">#29604</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofo…

  11. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11279

    <details open=""> <p>add GLM-5.3-Flash (GLM5-Next) support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27773">#27773</a>)</p> <ul> <li> <p>Rebase GLM-Next support onto master, and migrate to llama-memory-hybrid-idx</p> </li> <li> <p>Add i…

  12. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11278

    <details open=""> <p>vendor: update BoringSSL to 0.20260929.0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29669">#29669</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama…

  13. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11276

    <details open=""> <p>Hexagon f16 activation ops (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29209">#29209</a>)</p> <ul> <li>hexagon: add F16 support for activation ops (SILU/GELU/GELU_QUICK/GEGLU/SWIGLU)</li> </ul> <p>Widens ggml_hexagon_…

  14. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11275

    <details open=""> <p>gguf : reject tensor size that wraps after padding (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/26979">#26979</a>)</p> <p>GGML_PAD(nbytes, alignment) wraps to 0 when nbytes is within<br /> (alignment - 1) of SIZE_MAX, …

  15. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11274

    <details open=""> <p>ggml : check row bounds in get_rows_back (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29575">#29575</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </details>…

  16. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11273

    <details open=""> <p>model : support classifier_pooling for rerankers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29627">#29627</a>)</p> <ul> <li>model : support classifier_pooling for ModernBERT rerankers</li> </ul> <p>Assisted-by: Claud…

  17. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11272

    <details open=""> <p>hexagon: optimize concat op (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29673">#29673</a>)</p> <ul> <li>hex-concat: reduce pkts in gather/transpose hot loop</li> </ul> <p>gather directly into dst buffer, use special i…

  18. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11270

    <details open=""> <p>ggml-cuda: HIP: optimize packed byte subtraction (<code>__vsubss4</code> -&gt; <code>__vsub4</code>) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29478">#29478</a>)</p> <ul> <li> <p>ggml-cuda: HIP: optimize non-saturat…

  19. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11269

    <details open=""> <p>ci: add zdnn backend build but not test (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29541">#29541</a>)</p> <ul> <li>ci: add zdnn backend build but not test</li> </ul> <p>Signed-off-by: Aaron Teo <a href="mailto:aaron.…

  20. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11268

    <details open=""> <p>opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29555">#29555</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="no…

  21. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11267

    <details open=""> <p>vulkan: Tune GDN kernel, fix Intel performance (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29476">#29476</a>)</p> <ul> <li> <p>vulkan: tune GDN shader</p> </li> <li> <p>tune for Intel</p> </li> </ul> </details> <p><st…

  22. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11266

    <details open=""> <p>vulkan : Load F32 A matrix 2 at a time when its 2-aligned (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29254">#29254</a>)</p> <p>It turns out Intel doesn't particularly like loading F32s one at a<br /> time and we alre…

  23. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11265

    <details open=""> <p>vulkan: MOE aware mat_mul_id tile selection (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29182">#29182</a>)</p> <p>mut_mul_id selected its matmul tile with total token count.<br /> For MoE dispatch grid the true N per …

  24. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11264

    <details open=""> <p>ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29504">#29504</a>)</p> <ul> <li> <p>fix c++ odr by properly using GGML_COMMON_DECL_CPP</p> </li> <li> <p>using actu…

  25. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11263

    <details open=""> <p>vocab : keep NORMAL in PLaMo-2 and PLaMo-3 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29580">#29580</a>)</p> <ul> <li>vocab : keep NORMAL in PLaMo-2 and PLaMo-3</li> </ul> <p>The PLaMo-2 and PLaMo-3 vocabularies mark…

  26. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11262

    <details open=""> <p>ggml : accumulate f16 dot products in f32 on AVX512-FP16 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29545">#29545</a>)</p> <p>Supersedes <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp…

  27. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11261

    <details open=""> <p>ggml : require input tensors to be GGML_OP_NONE (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29647">#29647</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:…

  28. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11260

    <details open=""> <p>hexagon: add FP32 GELU_ERF and GEGLU_ERF support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29631">#29631</a>)</p> <ul> <li> <p>hexagon: add FP32 GELU_ERF and GEGLU_ERF support</p> </li> <li> <p>hex-erf: reduce regis…

  29. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11259

    <details open=""> <p>common : stop accepting draft tokens at EOG (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29638">#29638</a>)</p> <ul> <li> <p>common : stop accepting draft tokens at EOG</p> </li> <li> <p>cont : remove the test</p> </li…

  30. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11258

    <details open=""> <p>server : remove the built-in UI's service worker when the UI is not served (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29565">#29565</a>)</p> <p>With --path or --no-ui, /sw.js returned 404, and a 404 does not remove a…

  31. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11257

    <details open=""> <p>common : use fs::path for config dir (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29649">#29649</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </details> <p>…

  32. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11256

    <details open=""> <p>tests : adjust server string regex to also match m2 utlra results (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29648">#29648</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel…

  33. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11255

    <details open=""> <p>common : add fs_write_atomic() (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29642">#29642</a>)</p> <ul> <li>Check for buffered write errors when closing downloaded files.</li> <li>Use UTF-8 paths when writing ETag file…

  34. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11254

    <details open=""> <p>ggml : collect all input tensors into graph_inputs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29634">#29634</a>)</p> <p>graph_inputs was populated while splitting the graph, so it only<br /> contained the inputs that…

  35. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11249

    <details open=""> <p>llama : fix init in several tools/examples (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29632">#29632</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://lla…

  36. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11247

    <details open=""> <p>vulkan : reuse descriptor sets when bindings are constant (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29280">#29280</a>)</p> <ul> <li> <p>vulkan : reuse descriptor sets when bindings are constant</p> </li> <li> <p>vul…

  37. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11246

    <details open=""> <p>chat : fix Muse Glimmer ignoring response_format json_schema with --jinja (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29615">#29615</a>)</p> <ul> <li>chat : fix Muse Glimmer ignoring response_format json_schema with -…

  38. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11245

    <details open=""> <p>common : use fs::path for cache dirs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29595">#29595</a>)</p> <ul> <li>Avoid useless string conversions on Windows.</li> <li>No need for BSD or emscripten special cases.</li> …

  39. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11240

    <details open=""> <p>server : support typed content (vision/audio/video) input for /v1/embeddings endpoint (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29556">#29556</a>)</p> <ul> <li>server : support multimodal input for /v1/embeddings (Q…

  40. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11238

    <details open=""> <p>models: pad on the left with ggml_pad_ext (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29567">#29567</a>)</p> <ul> <li>models: pad on the left with ggml_pad_ext</li> </ul> <p>The Parakeet, LFM2-Audio, Granite Speech an…

  41. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11237

    <details open=""> <p>ggml-openvino: mark unaligned batch-stride views unsupported (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29603">#29603</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nof…

  42. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11236

    <details open=""> <p>batch: migrate speculative, mtmd and server to batch_ext (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29385">#29385</a>)</p> <ul> <li> <p>adapt common</p> </li> <li> <p>add common_batch</p> </li> <li> <p>wip</p> </li> …

  43. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11235

    <details open=""> <p>common : fix HF cache paths on Windows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29475">#29475</a>)</p> <p>Supersedes <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29158">#2915…

  44. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11234

    <details open=""> <p>webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29471">#29471</a>)</p> <ul> <li> <p>Fix: Handle unaligned writes in ggml_backend_webgpu_buffer_set_t…

  45. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11233

    <details open=""> <p>tests : refactor test-recurrent-state-rollback (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29426">#29426</a>)</p> <ul> <li>tests : use llama_context_ptr in test-recurrent-state-rollback</li> </ul> <p>Replace raw llama…

  46. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11232

    <details open=""> <p>ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29423">#29423</a>)</p> <ul> <li> <p>ggml-cpu: enable tiled flash attention for non-vector-mul…

  47. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11229

    <details open=""> <p>HIP: fix template skip for DKQ &gt; 256 mfma kernels (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29559">#29559</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">h…