PulseAugur
中
实时 11:09:39

llama.cpp 发布包括服务器修复、OpenVINO 优化和平台支持扩展

llama.cpp 项目发布了多个更新,包括 b11379、b11378、b11377、b11376、b11375、b11374、b11372、b11371 和 b11370 版本。这些版本在不同平台和功能上引入了各种改进和修复。值得注意的变化包括服务器端中止修复、JSON schema 聊天解析增强、OpenVINO 优化以及 CUDA 和 Metal 后端更新。该项目还继续扩展对 Snapdragon 和 openEuler 等不同硬件架构和操作系统的支持。 AI

影响 持续开发和优化本地 LLM 推理工具。

排序理由 这是一系列特定项目的软件发布说明,而非重大的行业事件。

在 llama.cpp — Releases 阅读 →

AI 生成摘要 · Google Gemini · 来自 73 个来源。 我们如何撰写摘要 →

llama.cpp 发布包括服务器修复、OpenVINO 优化和平台支持扩展

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是一系列特定项目的软件发布说明,而非重大的行业事件。
Source corroboration
73 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
7 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+17 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [73]

  1. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11399

    <details open=""> <p>CUDA: refactor swizzling code (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29612">#29612</a>)</p> <ul> <li> <p>CUDA: refactor swizzling code</p> </li> <li> <p>fix templates/loop bounds</p> </li> </ul> </details> <p><st…

  2. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11398

    <details open=""> <p>ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29806">#29806</a>)</p> <ul> <li> <p>ggml-cpu: vectorize BF16 K tails in tinyBLAS</p> </li> <li> <p>tests: Skip ti…

  3. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11397

    <details open=""> <p>cuda : move neu_padded to where it is used (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29940">#29940</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </detail…

  4. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11396

    <details open=""> <p>ci : windows llvm build requires ninja multi-config (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29959">#29959</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">ht…

  5. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11393

    <details open=""> <p>chat-peg-parser : clear current_tool when pending_tool_call is reset (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29942">#29942</a>)</p> <p>A TOOL_ID node that arrives after TOOL_CLOSE wrote through <code>current_tool<…

  6. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11392

    <details open=""> <p>ci : set default permissions (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29945">#29945</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li…

  7. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11391

    <details open=""> <p>cuda : move blocks_per_col to where it is used (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29939">#29939</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </de…

  8. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11390

    <details open=""> <p>CUDA: fix MMQ memory fault if n_expert &gt;&gt; n_ubatch (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29941">#29941</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollo…

  9. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11389

    <details open=""> <p>vulkan: fix rdna4 mat_vec tuning (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29934">#29934</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a>…

  10. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11388

    <details open=""> <p>imatrix: calculate activation-based statistics for new format (GGUF) imatrices (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/14891">#14891</a>)</p> <ul> <li>Use activations to calculate the stats</li> <li>Determine calc…

  11. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11387

    <details open=""> <p>spec : fix n-gram drafts rejected at temp &gt; 0 after truncation (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29924">#29924</a>)</p> <p>Co-authored-by: Pranesh Gonegandla <a href="mailto:[email protected]">pgoneg…

  12. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11386

    <details open=""> <p>common : prepare load_from_models_dir() for path conversion (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29674">#29674</a>)</p> <p>This is part of the fs::path modernization series.<br /> That was also the opportunity …

  13. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11385

    <details open=""> <p>server : fix dead LLAMA_ARG_HF_REPO_FILE key in preset allow-list (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29938">#29938</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">angt@huggingf…

  14. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11384

    <details open=""> <p>ci : pushing tag needs deploy key (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29937">#29937</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…

  15. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11382

    <details open=""> <p>webgpu: add f16 support to fill/set_rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29897">#29897</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama…

  16. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11381

    <details open=""> <p>mtmd : fix deprecated strdup warning on Windows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29863">#29863</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </d…

  17. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11380

    <details open=""> <p>vendor : update cpp-httplib to 0.59.0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29886">#29886</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.ap…

  18. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11379

    <details open=""> <p>server : fix laya abort by limiting n_batch to n_ubatch (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29903">#29903</a>)</p> <ul> <li>server : fix laya abort by limiting n_batch to n_ubatch</li> </ul> <p>Fixes <a class=…

  19. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11378

    <details open=""> <p>common : add common_is_tty() helper and fix deprecated warnings on Windows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29860">#29860</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">angt…

  20. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11377

    <details open=""> <p>chat : honor json_schema in Ling 3.0 parser (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29813">#29813</a>)</p> <ul> <li>chat : honor json_schema in Ling 3.0 parser</li> </ul> <p>Ling 3.0 only built a grammar for tool …

  21. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11376

    <details open=""> <p>ci: fix flaky ADD_ADD f16 by using the fused ADD tolerance (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29904">#29904</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofol…

  22. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11375

    <details open=""> <p>graph: gather the recurrent states once so the reserve covers every split (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29856">#29856</a>)</p> <p>build_rs gathered the extra states (n_rs - n_seqs rows) with their own<br…

  23. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11374

    <details open=""> <p>ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29852">#29852</a>)</p> <ul> <li>ggml-openvino : Qwen3.5 MoE perf (<a class="issu…

  24. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11372

    <details open=""> <p>qwen4exp : halve the indexer score memory (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29825">#29825</a>)</p> <ul> <li>qwen4exp : halve the indexer score memory</li> </ul> <p>The indexer scored all heads in one product…

  25. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11371

    <details open=""> <p>model: add support for clef decision model (text-only) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29831">#29831</a>)</p> <ul> <li> <p>init support for clef (text only)</p> </li> <li> <p>more static graph</p> </li> <l…

  26. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11370

    <details open=""> <p>CUDA: fuse shared experts into MMVQ (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29184">#29184</a>)</p> <ul> <li> <p>CUDA: fuse shared experts into MMVQ</p> </li> <li> <p>check if buffer is null</p> </li> <li> <p>move …

  27. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11368

    <details open=""> <p>spec : add probabilistic sampling for simple draft and MTP (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27694">#27694</a>)</p> <ul> <li> <p>Make the drafter probabilistic and the target verify by rejection sampling</p>…

  28. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11366

    <details open=""> <p>ggml-quants : avoid invalid rounding in qkx3 scale search (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29817">#29817</a>)</p> <ul> <li>ggml-quants : avoid invalid rounding in qkx3 scale search</li> </ul> <p>The imatrix…

  29. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11365

    <details open=""> <p>ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27096">#27096</a>)</p> <ul> <li>ggml-cpu : fix soft_max_back wrong output when dst aliases src1</li> </ul> <p…

  30. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11364

    <details open=""> <p>model: support nimble decision model (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29844">#29844</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app…

  31. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11362

    <details open=""> <p>metal : add tensor API flash attention kernel for F16 KV (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29570">#29570</a>)</p> <ul> <li> <p>metal : add tensor API flash attention kernel for F16 KV</p> </li> <li> <p>cont …

  32. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11361

    <details open=""> <p>llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29818">#29818</a>)</p> <ul> <li> <p>init conversion</p> </li> <li> <p>convert: ok</p> </li> <…

  33. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11355

    <details open=""> <p>vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28531">#28531</a>)</p> <p>Assisted-by: Claude Opus</p> </details> <p><strong>Website:</strong></p> …

  34. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11352

    <details open=""> <p>qwen4exp : optimize mask constructions (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29824">#29824</a>)</p> <ul> <li> <p>qwen4exp : optimize mask constructions</p> </li> <li> <p>cont : apply the same change for GLM5-nex…

  35. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11351

    <details open=""> <p>ggml : add <code>alloc_buffer_n</code> to buffer type interface (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23671">#23671</a>)</p> <ul> <li>ggml : add <code>alloc_buffer_n</code> to buffer type interface</li> </ul> <p…

  36. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11349

    <details open=""> <p>vulkan: add logging to pipeline compile issues (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29794">#29794</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:/…

  37. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11346

    <details open=""> <p>qwen4exp: fix tests (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29819">#29819</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a></li> </ul> <…

  38. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11345

    <details open=""> <p>hexagon: add q2_k and q3_k quant type support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29717">#29717</a>)</p> <ul> <li> <p>hexagon: add q2_k and q3_k quant type support</p> </li> <li> <p>hex-qk: consistent allocati…

  39. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11344

    <details open=""> <p>CUDA: fix 2 broken Volta FA cases (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29803">#29803</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.app</a…

  40. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11342

    <details open=""> <p>common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29816">#29816</a>)</p> <p>See <a href="https://gcc.gnu.org/bugzilla/show_bug.cgi?id=101510" rel="nofo…

  41. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11339

    <details open=""> <p>llama : clamp kpool re-pool bound to existing pools (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29805">#29805</a>)</p> <ul> <li> <p>tests : simplify function signature</p> </li> <li> <p>llama : clamp kpool re-pool bou…

  42. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11338

    <details open=""> <p>hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29685">#29685</a>)</p> <ul> <li> <p>hexagon: shared strided DMA copy for CPY and CONCAT, any-dim …

  43. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11337

    <details open=""> <p>server: return HTTP 400 for invalid embedding requests (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29060">#29060</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow"…

  44. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11335

    <details open=""> <p>cuda : route sm70 to the Turing MMVQ nwarps table (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29753">#29753</a>)</p> <ul> <li>cuda : route sm70 to the Turing MMVQ nwarps table</li> </ul> <p>Volta (sm_70) has no MMVQ p…

  45. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11334

    <details open=""> <p>metal : release temporary private transfer buffers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29777">#29777</a>)</p> <ul> <li>metal : release temporary private transfer buffers</li> </ul> <p>Assisted-by: OpenAI Codex…

  46. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11333

    <details open=""> <p>webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29358">#29358</a> (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29358">#…

  47. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11332

    <details open=""> <p>llama : fix invalid assert in recurrent memory (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29799">#29799</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https:/…

  48. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11331

    <details open=""> <p>CUDA: Handle compute type for NVFP4 on cublass path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29173">#29173</a>)</p> <ul> <li>CUDA: Handle compute type for NVFP4 on cublass path</li> </ul> <p>Signed-off-by: ynankani…

  49. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11330

    <details open=""> <p>Qwen4Exp: add MTP (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29761">#29761</a>)</p> <ul> <li> <p>Qwen4Exp: add MTP</p> </li> <li> <p>remove has_state member, check via ctx_bufs being non-empty</p> </li> <li> <p>consi…

  50. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11327

    <details open=""> <p>mtmd: cap max_image to ubatch for non_causal models (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29773">#29773</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">ht…

  51. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11326

    <details open=""> <p>meta: clear inactive AllReduce shards with FILL, not SCALE (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29793">#29793</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofol…

  52. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11325

    <details open=""> <p>jinja : skip copying loop scope unless a loop filter needs it (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29776">#29776</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="no…

  53. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11324

    <details open=""> <p>llama-mmap : avoid a second full-size copy of each tensor with direct-io (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29749">#29749</a>)</p> <p>Assisted-by: Claude</p> <p>Co-authored-by: Pranesh Gonegandla <a href="mai…

  54. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11323

    <details open=""> <p>HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29572">#29572</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llam…

  55. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11322

    <details open=""> <p>hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29785">#29785</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https:…

  56. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11321

    <details open=""> <p>BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29640">#29640</a>)</p> <ul> <li> <p>BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS</p> </li> <…

  57. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11320

    <details open=""> <p>common : add LLM-jp-4.1 Harmony dialect handler (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29681">#29681</a>)</p> <p>LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space<br /> after every special tok…

  58. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11319

    <details open=""> <p>opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29698">#29698</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" …

  59. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11318

    <details open=""> <p>vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29734">#29734</a>)</p> <ul> <li>vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3</li> </ul> <p>The original toke…

  60. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11317

    <details open=""> <p>llama-bench : fix verbosity filter to show GGML_LOG_ERROR (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28229">#28229</a>)</p> <ul> <li> <p>bench : fix verbosity filter to show GGML_LOG_ERROR (<a class="issue-link js-is…

  61. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11316

    <details open=""> <p>metal : use bf16 math for mxfp4 mul-mat (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29770">#29770</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.…

  62. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11313

    <details open=""> <p>model : re-enable -sm tensor for qwen4exp (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28569">#28569</a>)</p> <p><a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/27941">#27941</a> di…

  63. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11312

    <details open=""> <p>webgpu: fix SSM_SCAN binding aliasing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29750">#29750</a>)</p> </details> <p><strong>Website:</strong></p> <ul> <li><a href="https://llama.app" rel="nofollow">https://llama.ap…

  64. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11311

    <details open=""> <p>ggml-opencl : replace alloca() with std::vector (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29765">#29765</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </d…

  65. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11310

    <details open=""> <p>cuda: guard the iq4_nl dequantize row kernel against short rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29683">#29683</a>)</p> <p>dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than…

  66. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11309

    <details open=""> <p>Hexagon: optimize ALLREDUCE with support for safe scatter mode (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29757">#29757</a>)</p> <ul> <li> <p>hex-allreduce: add support for safe scatter mode</p> </li> <li> <p>hex-all…

  67. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11308

    <details open=""> <p>args: fix cli download mmproj arg (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28977">#28977</a>)</p> <ul> <li> <p>tests: add tests for cli download arg parsing</p> </li> <li> <p>args: fix cli download mmproj arg</p> <…

  68. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11307

    <details open=""> <p>llama : preserve original batch order for speculative decoding layer inputs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29019">#29019</a>)</p> <ul> <li>llama: preserve original batch order for layer inputs</li> </ul> …

  69. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11306

    <details open=""> <p>test-llama-archs : toggle causal_attn to catch graph shape changes (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29724">#29724</a>)</p> <p>After the device decode, flip causal_attn off, decode n_ubatch/2 then<br /> n_ub…

  70. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11304

    <details open=""> <p>llama: properly handle KV on training (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/28520">#28520</a>)</p> <ul> <li> <p>llama: properly handle KV on training</p> </li> <li> <p>improve</p> </li> </ul> </details> <p><stro…

  71. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11303

    <details open=""> <p>batch: migrate the rest of examples to llama_batch_ext (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29601">#29601</a>)</p> <ul> <li> <p>migrate the rest</p> </li> <li> <p>test-thread-safety</p> </li> <li> <p>rm common_…

  72. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11302

    <details open=""> <p>glm5-next: give dead indexer slots unique scatter rows (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29745">#29745</a>)</p> <p>The sparse indexer mask is built with a set_rows scatter. Padded pools,<br /> absent sequenc…

  73. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b11301

    <details open=""> <p>ggml/gguf : fix integer overflow (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/29384">#29384</a>)</p> <ul> <li> <p>ggml: fix integer overflow guard for zero-element tensors</p> </li> <li> <p>ggml: validate number of ele…