PulseAugur
EN
LIVE 13:16:40

llama.cpp releases multiple updates with performance and bug fixes

The llama.cpp project has released several updates, including versions b9975, b9974, b9973, b9972, b9971, b9970, b9969, b9968, b9966, and b9965. These releases introduce various improvements and bug fixes across multiple platforms such as macOS, Linux, Android, and Windows. Notable changes include optimizations for DeepSeek models, enhancements to Vulkan and OpenCL performance on specific hardware like Adreno GPUs, and fixes for CUDA memory querying issues. AI

IMPACT These updates enhance the performance and stability of the llama.cpp project, potentially improving the efficiency of running large language models on various hardware.

RANK_REASON The cluster consists of multiple minor release notes for the llama.cpp project, detailing bug fixes and performance improvements rather than a new model release or significant research.

Read on llama.cpp — Releases →

AI-generated summary · Google Gemini · from 150 sources. How we write summaries →

llama.cpp releases multiple updates with performance and bug fixes

COVERAGE [150]

  1. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9999

    <details open=""> <p>kleidiai : add SME2 f32 kernel (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24414">#24414</a>)</p> <ul> <li> <p>kleidiai : add SME2 f32 kernel</p> </li> <li> <p>enable dynamic scheduling for SME2 f32 kernel</p> </li> <…

  2. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9996

    <details open=""> <p>arg: Flush log before exiting after usage() (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25504">#25504</a>)</p> <p>Under certain conditions, it's possible for messages emitted via LOG()<br /> to get lost before exit, a…

  3. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9995

    <details open=""> <p>sycl: set fattn_vec_nthreads to 256 for Battlemage (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25205">#25205</a>)</p> <p>Currently detects lunarlake + battlemage / xe2 and<br /> sets the value to 256.</p> <p>Keeps def…

  4. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9994

    <details open=""> <p>metal : add Q2_0 support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25419">#25419</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download/b9994…

  5. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9993

    <details open=""> <p>model: add Hy3 (hy_v3) support with MTP speculative decoding (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25395">#25395</a>)</p> <ul> <li>model: add Hy3 (hy_v3) architecture support</li> </ul> <p>Adds Tencent Hunyuan 3…

  6. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9992

    <details open=""> <p>CUDA: refactor MMQ kernel configuration (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24127">#24127</a>)</p> <ul> <li> <p>CUDA: refactor MMQ kernel configuration</p> </li> <li> <p>fix Blackwell config</p> </li> <li> <p>…

  7. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9990

    <details open=""> <p>spec: add Minimax2 eagle3 support</p> <ul> <li> <p>Fix nullptr in minimax2 EAGLE3</p> </li> <li> <p>minor : add newline</p> </li> </ul> <hr /> <p>Co-authored-by: Georgi Gerganov <a href="mailto:[email protected]">[email protected]</a></p> </details> <p><s…

  8. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9988

    <details open=""> <p>tests: Harmonize header use (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25616">#25616</a>)</p> <ul> <li> <p>tests: Harmonize the use of private ggml includes</p> </li> <li> <p>tests: In test-backend-ops, use quoted in…

  9. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9987

    <details open=""> <p>gguf : add tensor shape accessor (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24405">#24405</a>)</p> <ul> <li> <p>gguf : add tensor shape accessors</p> </li> <li> <p>gguf : return tensor shape as const int64_t *</p> </…

  10. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9986

    <details open=""> <p>chat : fix reasoning leak with force-opened bare templates (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24674">#24674</a>)</p> <ul> <li>chat : fix reasoning leak with force-opened bare templates</li> </ul> <p>The reaso…

  11. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9985

    <details open=""> <p>sycl: add fused top-k MoE (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25217">#25217</a>)</p> <ul> <li> <p>sycl: add fused top-k MoE</p> </li> <li> <p>sycl: address review: GGML_SYCL_ENABLE_FUSION env, move fusion disp…

  12. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9984

    <details open=""> <p>sycl: add Q2_K to DMMV reorder path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25064">#25064</a>)</p> <p>Signed-off-by: Todd Malsbary <a href="mailto:[email protected]">[email protected]</a></p> </details…

  13. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9982

    <details open=""> <p>server: honour per-request reasoning_budget_tokens in chat completions (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23116">#23116</a>)</p> <ul> <li>server: honour per-request reasoning_budget_tokens in chat completions…

  14. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9981

    <details open=""> <p>vendor : update cpp-httplib to 0.50.1 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25576">#25576</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/d…

  15. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9980

    <details open=""> <p>server: Don't consider models with --no-mmproj-auto as multimodal (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25590">#25590</a>)</p> <p>If mmproj is explicitly disabled via the model preset or command-line<br /> param…

  16. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9979

    <details open=""> <p>mtmd: fix silent prompt truncation on embedded NUL (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25548">#25548</a>)</p> <ul> <li>mtmd: fix silent prompt truncation on embedded NUL</li> </ul> <p>mtmd_input_text carried t…

  17. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9978

    <details open=""> <p>server : evict checkpoints within min-step of each other (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25472">#25472</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/l…

  18. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9977

    <details open=""> <p>server : fix image blocks in tool_result being dropped during Anthropic OpenAI conversion (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/22536">#22536</a>)</p> <ul> <li>server : fix image blocks in tool_result being drop…

  19. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9975

    <details open=""> <p>gguf : reject empty metadata keys (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24917">#24917</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/downl…

  20. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9974

    <details open=""> <p>cuda: Don't crash when querying memory on device with no free memory. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25157">#25157</a>)</p> <p>If a Cuda device has no or limited available memory, the actual call<br /> to…

  21. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9973

    <details open=""> <p>DeepseekV4: clear cache only for seq rather than full (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25521">#25521</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llam…

  22. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9972

    <details open=""> <p>server: allow stream for exec_shell_command (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25526">#25526</a>)</p> <ul> <li> <p>init stream</p> </li> <li> <p>add stream for shell tool</p> </li> <li> <p>add test</p> </li> …

  23. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9971

    <details open=""> <p>server: refactor server_stream (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25541">#25541</a>)</p> <ul> <li> <p>server: refactoring, remove spipe from server_http_res</p> </li> <li> <p>wip</p> </li> <li> <p>remove non-…

  24. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9970

    <details open=""> <p>ggml : add GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24231">#24231</a>)</p> <ul> <li> <p>ggml : add GGML_OP_LIGHTNING_INDEXER that impleme…

  25. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9969

    <details open=""> <p>Vulkan: route large matmuls to medium tile on Adreno (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24877">#24877</a>)</p> <ul> <li>[Vulkan] Fixes llama-cli breaking over longer promts sizes</li> </ul> <p>The llama-cli w…

  26. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9968

    <details open=""> <p>opencl: add int8 dp4 dense and MoE prefill optimization for Adreno GPUs (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25537">#25537</a>)</p> <ul> <li> <p>opencl: add int8 dp4 dense and moe GEMM</p> </li> <li> <p>opencl:…

  27. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9966

    <details open=""> <p>llama : make tensor-split regex patterns static (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24710">#24710</a>)</p> <p>llama_meta_device_get_split_state() recompiled 29 std::regex on every call.<br /> In -sm tensor mod…

  28. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9965

    <details open=""> <p>hexagon: improve ARGSORT performance for small tensors (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25512">#25512</a>)</p> <ul> <li> <p>hex-sort: add efficient bitomic sort in hvx regs up to 1024 elements</p> </li> <li…

  29. llama.cpp — Releases TIER_1 (SO) · kdkd ·

    b9976

    <p>Fix conditional to display 'LLAMA_SPLIT_MODE_TENSOR not implemented f…</p>

  30. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9964

    <details open=""> <p>arg: prevent duplicate spec model downloads (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25527">#25527</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/rele…

  31. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9963

    <details open=""> <p>mtmd: deepseek-ocr v1 multi-tile (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24717">#24717</a>)</p> <ul> <li> <p>mtmd: deepseek-ocr v1 multi-tile dynamic resolution + unified image-preprocessors for both versions (ds-…

  32. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9960

    <details open=""> <p>server: remove loading.html (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25500">#25500</a>)</p> <ul> <li> <p>server: remove loading.html</p> </li> <li> <p>apply ui changes</p> </li> </ul> </details> <p><strong>macOS/iO…

  33. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9959

    <details open=""> <p>sync : ggml</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download/b9959/llama-b9959-bin-macos-arm64.tar.gz">macOS Apple Silicon (arm64)</a></li> <li>macOS Apple Silicon (arm64, KleidiAI ena…

  34. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9957

    <details open=""> <p>server: improve tools, remove apply_diff (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25498">#25498</a>)</p> <ul> <li> <p>server: improve tools, remove apply_diff</p> </li> <li> <p>improve edit tool</p> </li> <li> <p>a…

  35. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9956

    <details open=""> <p>cli: fix crash on wrong server base url (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25497">#25497</a>)</p> <ul> <li> <p>llama-cli: fix crash on wrong server base url by catching exceptions and graceful exit</p> </li> …

  36. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9952

    <details open=""> <p>llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25370">#25370</a>)</p> <ul> <li> <p>llama : make all KQ masks (e…

  37. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9950

    <details open=""> <p>llama-batch: add unit test (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25471">#25471</a>)</p> <ul> <li> <p>llama-batch: add unit test</p> </li> <li> <p>fix win32 builds</p> </li> <li> <p>add not implemented assertion …

  38. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9949

    <details open=""> <p>opencl: cluster-parallel decode FA for Adreno (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25473">#25473</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/re…

  39. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9948

    <details open=""> <p>ggml : process data in smaller chunks in CUDA ggml_top_k() and ggml_argsort() to reduce temporary buffers memory usage (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24776">#24776</a>)</p> <ul> <li> <p>ggml : process dat…

  40. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9947

    <details open=""> <p>cli: add --output option (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25484">#25484</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download/b9947…

  41. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9946

    <details open=""> <p>hexagon: tiling, tracing and optimizations for unary ops (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25474">#25474</a>)</p> <ul> <li> <p>hexagon: tile wide rows in pointwise unary ops to avoid VTCM overflow</p> </li> …

  42. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9945

    <details open=""> <p>server : move chat-template thinking probe inside the init try/catch (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24093">#24093</a>)</p> <p>A model whose chat template parses at init but fails parser generation<br /> a…

  43. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9941

    <details open=""> <p>Only index by compile times + always multiply/add (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25445">#25445</a>)</p> <p>The first one avoids relying on compile to optimize local memory away,<br /> and the second is ch…

  44. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9940

    <details open=""> <p>llama-bench : init params.offline (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25476">#25476</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </details> <p><st…

  45. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9939

    <details open=""> <p>metal : add CONV_2D_DW (depthwise convolution) support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/21565">#21565</a>)</p> <ul> <li> <p>metal : add CONV_2D_DW (depthwise 2D convolution) support</p> </li> <li> <p>test :…

  46. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9938

    <details open=""> <p>ggml-hip: enable -funsafe-math-optimizations (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24668">#24668</a>)</p> <p>CUDA is compiled with fast math and AMD/HIP is not — this flag lets AMD use fast math too.</p> <p>We c…

  47. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9937

    <details open=""> <p>cuda: align snake fusion matcher with the other backends (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25460">#25460</a>)</p> <ul> <li>cuda: fix snake fusion type predicate, a and inv_b are F32</li> </ul> <p>The matcher…

  48. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9936

    <details open=""> <p>server : respect min-step when splitting prompt batches (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25420">#25420</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/ll…

  49. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9935

    <details open=""> <p>hexagon: add VISION RoPE support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25216">#25216</a>)</p> <ul> <li> <p>hexagon: add VISION RoPE support</p> </li> <li> <p>hexagon: support RoPE on strided half-dim views for a…

  50. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9934

    <details open=""> <p>ggml-webgpu: tune subgroup split (d_split) in flash_attn_vec (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25418">#25418</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-o…

  51. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9933

    <details open=""> <p>opencl: Q6_K GEMM/GEMV fix for ne01 of weights that are not multiples of 128. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25464">#25464</a>)</p> <ul> <li>opencl: fix garbled output for Q6_K weights with ne01 % 128 != …

  52. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9932

    <details open=""> <p>vulkan: disable FA mask_opt on GCN to improve performance (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24362">#24362</a>)</p> <ul> <li> <p>vulkan: disable FA mask_opt on GCN to improve performance</p> </li> <li> <p>ree…

  53. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9931

    <details open=""> <p>opencl: ragged-tile MoE prefill FP16 GEMM optimization (skip padded expert tiles) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25433">#25433</a>)</p> <ul> <li>opencl: ragged-tile MoE prefill GEMM (skip padded expert ti…

  54. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9930

    <details open=""> <p>llama-batch: fix allowed decreasing pos in a seq (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25449">#25449</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp…

  55. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9929

    <details open=""> <p>vulkan: for small AMD GPUs, reduce submission threshold based on CU count (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25240">#25240</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://gith…

  56. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9928

    <details open=""> <p>hexagon: new vtcm layouts and improved pipelines for MUL_MAT, MUL_MAT_ID and FLASH_ATTN_EXT (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25425">#25425</a>)</p> <ul> <li> <p>hex-fa: refactor kernel param compute to use …

  57. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9927

    <details open=""> <p>cli : move to HTTP-based implementation (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24948">#24948</a>)</p> <ul> <li> <p>cli: move to HTTP-based implementation</p> </li> <li> <p>wip</p> </li> <li> <p>working</p> </li> …

  58. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9925

    <details open=""> <p>cuda : add support for f16-&gt;f16 GGML_OP_SET_ROWS (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25367">#25367</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.…

  59. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9924

    <details open=""> <p>llama: refactor fused ops (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24646">#24646</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download/b992…

  60. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9923

    <details open=""> <p>server-stream: follow-up on SSE Replay Buffer (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23226">#23226</a>) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25047">#25047</a>)</p>…

  61. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9922

    <details open=""> <p>llama-batch: add n_keep_tail in split_equal for recurrent models (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25278">#25278</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/gg…

  62. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9918

    <details open=""> <p>metal : add set_rows with src0 f16 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25434">#25434</a>)</p> <p>Co-authored-by: Georgi Gerganov <a href="mailto:[email protected]">[email protected]</a></p> </details> <p><…

  63. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9916

    <details open=""> <p>ggml : fix A indexing in simd_gemm scalar tail-column path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25390">#25390</a>)</p> <p><code>simd_gemm()</code> has an incorrect A-matrix index in the scalar tail-column path …

  64. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9915

    <details open=""> <p>ggml : add support for CPU f16-&gt;f16 GGML_OP_SET_ROWS (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25344">#25344</a>)</p> <ul> <li> <p>ggml : add support for CPU f16-&gt;f16 GGML_OP_SET_ROWS</p> </li> <li> <p>ggml : …

  65. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9914

    <details open=""> <p>opencl: fix potential crash in aos reconstruct (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25383">#25383</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/r…

  66. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9913

    <details open=""> <p>Add Q2_0 quantization: type definition and CPU backend (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24448">#24448</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/lla…

  67. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9912

    <details open=""> <p>spec : fix naming, spacing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25410">#25410</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download/b99…

  68. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9911

    <details open=""> <p>CUDA: Fuse MMVQ post-scale for NVFP4 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24481">#24481</a>)</p> <ul> <li>CUDA: Fuse MMVQ for NVFP4 and BS 1</li> </ul> <p>TODO:</p> <ol> <li>Add tests to test-backend-ops (did v…

  69. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9910

    <details open=""> <p>server : fix draft model fit vs load inconsistency (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25056">#25056</a>)</p> <ul> <li> <p>fix: draft model fit vs load inconsistency</p> </li> <li> <p>refactor(server): unify d…

  70. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9909

    <details open=""> <p>server : add timings and progress to /responses API stream (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25348">#25348</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org…

  71. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9908

    <details open=""> <p>server: enforce prompt cache RAM limit (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25070">#25070</a>)</p> <p>Before this commit, --cache-ram was not a hard limit:</p> <ul> <li>The cache always kept at least one entry,…

  72. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9907

    <details open=""> <p>common : add missing include in common.h (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25220">#25220</a>)</p> <p>Signed-off-by: zhangrunda <a href="mailto:[email protected]">[email protected]</a></p> <…

  73. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9906

    <details open=""> <p>ggml-hip : add -fno-finite-math-only alongside -ffast-math (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25373">#25373</a>)</p> <p>-ffast-math implies -ffinite-math-only under ROCm/clang 22, which<br /> disables INFINIT…

  74. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9905

    <details open=""> <p>llama: fix quantized kv-cache for dsv4 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25202">#25202</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/…

  75. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9904

    <details open=""> <p>[SYCL] fix unsupported UT cases of CONT &amp; CPY (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25231">#25231</a>)</p> <ul> <li> <p>fix unsupported UT cases of CONT &amp; CPY</p> </li> <li> <p>update ops.md</p> </li> <l…

  76. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9902

    <details open=""> <p>[SYCL] support OP cross_entropy_loss, cross_entropy_loss_back (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25236">#25236</a>)</p> <ul> <li> <p>support OP cross_entropy_loss, cross_entropy_loss_back</p> </li> <li> <p>co…

  77. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9901

    <details open=""> <p>sycl : set K_QUANTS_PER_ITERATION to 1 on DMMV path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25063">#25063</a>)</p> <ul> <li>sycl: add supported types to ggml_sycl_supports_reorder_dmmv</li> </ul> <p>The reordered …

  78. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9899

    <details open=""> <p>sycl : enhance argsort to support all UT cases (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25125">#25125</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/r…

  79. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9898

    <details open=""> <p>sycl : use sycl func to fix AOT double type issue (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25081">#25081</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cp…

  80. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9897

    <details open=""> <p>sycl : rename the env vars from "disable" to "enable" (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25042">#25042</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llam…

  81. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9895

    <details open=""> <p>speculative : fix out-of-bounds read in ngram-map on prompt shrink (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23936">#23936</a>)</p> <ul> <li> <p>speculative : fix out-of-bounds read in ngram-map on prompt shrink</p>…

  82. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9894

    <details open=""> <p>vulkan : check src0 type in GGML_OP_SET_ROWS to avoid failures due to unimplemented f16 support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25351">#25351</a>)</p> <ul> <li> <p>vulkan : check src0 type in GGML_OP_SET_R…

  83. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9893

    <details open=""> <p>opencl: general flash attention decode performance optimizations (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25366">#25366</a>)</p> <ul> <li> <p>opencl: vec flash-attention decode kernels for f16/q8_0/q4_0 KV</p> </li…

  84. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9892

    <details open=""> <p>common: Set optimal default thread count for ppc ( linux as well as AIX) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25237">#25237</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://githu…

  85. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9891

    <details open=""> <p>metal: add col2im_1d op (f32/f16/bf16) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25176">#25176</a>)</p> <ul> <li>metal: add col2im_1d op (f32/f16/bf16)</li> </ul> <p>Gather kernel mirroring the CPU/CUDA path: each o…

  86. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9890

    <details open=""> <p>CUDA: remove -sm row, refactor cuBLAS (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24216">#24216</a>)</p> <ul> <li> <p>CUDA: remove -sm row, refactor cuBLAS</p> </li> <li> <p>fix CDNA + BF16 logic</p> </li> <li> <p>fix…

  87. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9888

    <details open=""> <p>CUDA: extend K-type validation to V-types for flash attention (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24403">#24403</a>)</p> <ul> <li> <p>CUDA: extend K-type validation to V-types for flash attention</p> </li> <li…

  88. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9886

    <details open=""> <p>ggml-cpu: use UE4M3 LUT in ARM NVFP4 dot product (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25331">#25331</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp…

  89. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9885

    <details open=""> <p>ggml-cpu: Enable tiled matmul on AIX (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25199">#25199</a>)</p> <p>The matmul_tiled path uses large local stack buffers for A_pack and B_pack. On AIX this can trigger a segmenta…

  90. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9884

    <details open=""> <p>vulkan: fix 32-bit integer overflow in CEIL_DIV (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25245">#25245</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/…

  91. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9882

    <details open=""> <p>scripts : use HF_TOKEN when downloading UI assets (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25280">#25280</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> <…

  92. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9881

    <details open=""> <p>ggml-hip: enable -ffast-math for HIP builds (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23862">#23862</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/rele…

  93. llama.cpp — Releases TIER_1 (SO) · adavyas ·

    b9879

    <p>ggml-cuda: optimize conv_transpose_1d indexing (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25310">#25310</a>)</p>

  94. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9878

    <details open=""> <p>Fix stale tensor-split params for draft models (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24814">#24814</a>)</p> <ul> <li> <p>meta: fix tensor split metadata for GQA attention</p> </li> <li> <p>Tidied the code a bit …

  95. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9877

    <details open=""> <p>abort if we see a multi buffer (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25276">#25276</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download…

  96. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9876

    <details open=""> <p>ggml : fix tensor-parallel + -ncmoe crash on MoE models (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25028">#25028</a>)</p> <p>Tensor parallelism (-sm tensor) combined with -ncmoe (CPU-offloaded MoE<br /> experts) abor…

  97. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9874

    <details open=""> <p>cuda : concat implementation for quantized types (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25303">#25303</a>)</p> <ul> <li> <p>cuda : concat implementation for quantized types</p> </li> <li> <p>chore : apply am17an …

  98. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9873

    <details open=""> <p>llama : add guard for K/V rotation input when buffer is unallocated (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25215">#25215</a>)</p> <p>llm_graph_input_attn_kv::set_input and llm_graph_input_attn_kv_iswa::set_input<…

  99. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9871

    <details open=""> <p>ggml : fix broken CPU concat implementation for quantized types (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25247">#25247</a>)</p> <ul> <li> <p>ggml : fix broken CPU concat implementation for quantized types</p> </li>…

  100. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9870

    <details open=""> <p>chat: trim messages sent to StepFun parser (fixes long reasoning loops) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25238">#25238</a>)</p> <ul> <li> <p>chat: trim messages sent to StepFun parser (fixes long reasoning …

  101. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9867

    <details open=""> <p>spec: support spec-draft-p-min in DFlash (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25246">#25246</a>)</p> <ul> <li> <p>spec: support spec-draft-p-min in DFlash</p> </li> <li> <p>dflash: add n_min guard</p> </li> <li…

  102. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9866

    <details open=""> <p>cuda: enable topk-moe fusion for 288 experts (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25267">#25267</a>)</p> <ul> <li>cuda: enable topk-moe fusion for 288 experts</li> </ul> <p>The topk-moe fusion only accepted pow…

  103. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9864

    <details open=""> <p>server + ui: ping silent SSE streams every 1s and kick only after 3s so slow prefill never drops healthy connections (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25241">#25241</a>)</p> <ul> <li> <p>server + ui: ping si…

  104. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9862

    <details open=""> <p>Remove redundant CUDA copies after gated_delta_net. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23940">#23940</a>)</p> <ul> <li>Remove redundant CUDA copies after gated_delta_net.</li> </ul> <p>Currently, GDN writes r…

  105. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9861

    <details open=""> <p>vendor : update cpp-httplib to 0.49.0 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25218">#25218</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/d…

  106. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9860

    <details open=""> <p>llama : add llama_model_ftype_name() (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25134">#25134</a>)</p> <ul> <li>llama : add llama_model_ftype_name()</li> </ul> <p>Expose the model file type (quantization) name, e.g. …

  107. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9859

    <details open=""> <p>opencl: allow loading precompiled binary kernels from library (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23042">#23042</a>)</p> <ul> <li> <p>opencl: allow loading binary kernel</p> </li> <li> <p>opencl: add libdl.h</…

  108. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9858

    <details open=""> <p>common : use hf primary split as model path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25194">#25194</a>)</p> <p>Fixes <a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/issues/25181">#25…

  109. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9857

    <details open=""> <p>hexagon: flash attention rework (optimizations, accuracy improvements, etc) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25085">#25085</a>)</p> <ul> <li> <p>hex-mm: fold mm quant tasks into the main matmul threads</p> …

  110. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9856

    <details open=""> <p>CUDA: consistent use of <strong>restrict</strong> + PDL for FA (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25185">#25185</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml…

  111. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9855

    <details open=""> <p>ggml-cpu: add AVX2 optimization for nvfp4 dot product and use UE4M3 LUT (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23961">#23961</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github…

  112. llama.cpp — Releases TIER_1 (SO) · allozaur ·

    b9853

    <p>ui: Remove PWA navigate fallback to prevent caching API endpoint requ…</p>

  113. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9852

    <details open=""> <p>opencl: initial q1_0 support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25160">#25160</a>)</p> <ul> <li> <p>opencl: general q1_0 support</p> </li> <li> <p>opencl: add Adreno GEMM/GEMV for q1_0</p> </li> </ul> </detai…

  114. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9851

    <details open=""> <p>cuda : prevent integer truncation and overflow errors when using KQ mask strides in flash_attn_mask_to_KV_max kernel (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24945">#24945</a>)</p> <p>Co-authored-by: Stanisław Szym…

  115. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9850

    <details open=""> <p>model : register t_layer_inp for qwen3next (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25141">#25141</a>)</p> <ul> <li>Fix input assignment in layer processing loop</li> </ul> <p>Fix DFLASH for qwen-coder-next</p> <ul…

  116. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9849

    <details open=""> <p>common,server: handle bracketed IPv6 literals in URL authority (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25140">#25140</a>)</p> <ul> <li>common,server: handle bracketed IPv6 literals in URL authority</li> </ul> <p>P…

  117. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9848

    <details open=""> <p>CUDA: fix get_rows_back for tables with more than 65535 rows (grid-y clamp + stride) (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25103">#25103</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="h…

  118. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9847

    <details open=""> <p>CUDA: fix Gemma E4B MTP FlashAttention (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25148">#25148</a>)</p> <ul> <li> <p>CUDA: fix Gemma E4B MTP FlashAttention</p> </li> <li> <p>remove unused template declaration</p> </…

  119. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9846

    <details open=""> <p>vulkan: roll bk loop in matmul for asahi linux (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24663">#24663</a>)</p> <ul> <li> <p>vulkan: roll bk loop in matmul for asahi linux</p> </li> <li> <p>vulkan: fix inline commen…

  120. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9844

    <details open=""> <p>ggml-webgpu: add support for NVFP4 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25143">#25143</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/down…

  121. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9843

    <details open=""> <p>Revert "sched : reintroduce less synchronizations during split compute (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/20793">#20793</a>)" (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/p…

  122. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9842

    <details open=""> <p>common : dedup preset and cached model entries in /v1/models (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25131">#25131</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]

  123. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9840

    <details open=""> <p>DeepSeek V4 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24162">#24162</a>)</p> <ul> <li> <p>convert: add dsv4 conversion</p> </li> <li> <p>add basic setup</p> </li> <li> <p>add llm_graph_input_dsv4</p> </li> <li> <p>a…

  124. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9839

    <details open=""> <p>tools/ui: restore Tailwind scanning in ignored worktrees (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24879">#24879</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/l…

  125. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9838

    <details open=""> <p>common : remove unused regex-partial (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25118">#25118</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/do…

  126. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9837

    <details open=""> <p>jinja, chat: add --reasoning-preserve flag (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25105">#25105</a>)</p> <ul> <li> <p>jinja, chat: add --reasoning-preserve flag</p> </li> <li> <p>correct help message</p> </li> </…

  127. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9835

    <details open=""> <p>ui: fix stop and reasoning skip in single-model mode (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25084">#25084</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama…

  128. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9833

    <details open=""> <p>chat : implement minicpm5 parser (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24889">#24889</a>)</p> <ul> <li> <p>Add minicpm5 tool call parser</p> </li> <li> <p>Refactor MiniCPM5 PEG parser per review feedback</p> </l…

  129. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9832

    <details open=""> <p>jinja: add --dump-prog for debugging (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25086">#25086</a>)</p> <ul> <li> <p>jinja: add --dump-prog for debugging</p> </li> <li> <p>Update common/jinja/runtime.cpp</p> </li> </u…

  130. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9831

    <details open=""> <p>spec : add DFlash support (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/22105">#22105</a>)</p> <ul> <li> <p>spec: add DFlash v2 support</p> </li> <li> <p>dflash: support sliding window attention per layer_types</p> </li…

  131. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9830

    <details open=""> <p>common : allow --offline in llama download (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25091">#25091</a>)</p> <p>Expose the existing --offline flag to <code>llama download</code> so a script can<br /> run it to check …

  132. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9829

    <details open=""> <p>logs : reduce v2 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25078">#25078</a>)</p> <ul> <li> <p>server : reduce logs</p> </li> <li> <p>cont : common</p> </li> <li> <p>cont : spec</p> </li> <li> <p>cont : CMN_ -&gt; C…

  133. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9828

    <details open=""> <p>opencl: flash attention improvement (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25069">#25069</a>)</p> <ul> <li> <p>opencl: rework FA kernel for f16 and f32</p> </li> <li> <p>opencl: flash-attention prefill prepass ke…

  134. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9827

    <details open=""> <p>[CUDA] Added a cudaMemcpy2DAsync fast path to ggml_cuda_cpy (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25057">#25057</a>)</p> <ul> <li>[CUDA] Added a cudaMemcpy2DAsync fast path to ggml_cuda_cpy</li> </ul> <p>Add a C…

  135. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9826

    <details open=""> <p>sycl : fix failed ut cases of norm (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25044">#25044</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/down…

  136. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9825

    <details open=""> <p>vulkan: fix step operator for 0 input (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25036">#25036</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/d…

  137. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9824

    <details open=""> <p>binaries : Improve rpc-server and export-graph-ops names. (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25045">#25045</a>)</p> <p>Tests are generally prefixed with -test, so rename export-graph-ops<br /> accordingly.</p…

  138. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9823

    <details open=""> <p>ci : add windows-openvino to check-release (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25022">#25022</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/relea…

  139. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9822

    <details open=""> <p>tests : fix test-chat-template --no-common option (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25075">#25075</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cp…

  140. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9821

    <details open=""> <p>app : allow --version, --licenses &amp; --help (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25054">#25054</a>)</p> <p>Signed-off-by: Adrien Gallouët <a href="mailto:[email protected]">[email protected]</a></p> </de…

  141. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9820

    <details open=""> <p>sched : reintroduce less synchronizations during split compute (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/20793">#20793</a>)</p> <ul> <li> <p>CUDA: Improve performance via less synchronizations between token (<a clas…

  142. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9817

    <details open=""> <p>openvino: Update to OV 2026.2.1, self-contained release packages, operator improvements (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24974">#24974</a>)</p> <ul> <li> <p>Update to OV 2026.2.1, Make OV release packages s…

  143. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9816

    <details open=""> <p>sync : ggml</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download/b9816/llama-b9816-bin-macos-arm64.tar.gz">macOS Apple Silicon (arm64)</a></li> <li>macOS Apple Silicon (arm64, KleidiAI ena…

  144. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9814

    <details open=""> <p>vulkan: opt mul_mat_vecq for mi50 (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/22933">#22933</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/downl…

  145. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9813

    <details open=""> <p>vulkan: add INTEL_XE1 arch enum and enable coopmat1 on Intel Xe-LPG Plus (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24404">#24404</a>)</p> <ul> <li>vulkan: add INTEL_PRE_XE2 arch enum and enable coopmat1 on Intel Xe-…

  146. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9811

    <details open=""> <p>vulkan: Workaround compiler bug in conv2d coopmat2 path (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/24924">#24924</a>)</p> <ul> <li> <p>vulkan: Workaround compiler bug in conv2d coopmat2 path</p> </li> <li> <p>apply s…

  147. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9810

    <details open=""> <p>CUDA: add cublasSgemmBatched mapping for HIP/MUSA vendor headers (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25033">#25033</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/gg…

  148. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9804

    <details open=""> <p>mamba2: remove hardcoded 2x expansion factor and invalid d_inner % d_state check (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/23082">#23082</a>)</p> <ul> <li> <p>mamba2: remove hardcoded 2x expansion factor, support an…

  149. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9803

    <details open=""> <p>opencl: flush profiling batch at shutdown for incomplete batches (<a class="issue-link js-issue-link" href="https://github.com/ggml-org/llama.cpp/pull/25016">#25016</a>)</p> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/gg…

  150. llama.cpp — Releases TIER_1 (SO) · github-actions[bot] ·

    b9802

    <details open=""> </details> <p><strong>macOS/iOS:</strong></p> <ul> <li><a href="https://github.com/ggml-org/llama.cpp/releases/download/b9802/llama-b9802-bin-macos-arm64.tar.gz">macOS Apple Silicon (arm64)</a></li> <li>macOS Apple Silicon (arm64, KleidiAI enabled) <a href="http…