ggml-org
PulseAugur coverage of ggml-org — every cluster mentioning ggml-org across labs, papers, and developer communities, ranked by signal.
10 day(s) with sentiment data
llama.cpp to introduce more advanced KV cache optimizations
Given the recent mention of optimized KV cache for Qwen 3.8 27B and the general trend of memory reduction efforts in LLM inference, it is likely that llama.cpp will see further development in KV cache management techniques to support even longer contexts and larger models on consumer hardware.
llama.cpp actively expanding model support and hardware optimizations
Recent evidence shows multiple pull requests and updates to llama.cpp, including integration of new models like Kimi-K3 and performance optimizations such as AVX2 for prompt processing and CPU offload for dense models. This indicates a strong development velocity focused on broadening compatibility and enhancing efficiency on diverse hardware.
Increased adoption of GGUF format for new model releases
The availability of Qwen 3.8 27B in GGUF format and its integration into llama.cpp suggests a growing trend. We hypothesize that more model providers will release their models in GGUF or similar formats optimized for llama.cpp, increasing the accessibility of cutting-edge models for local inference.
-
llama.cpp adds support for K2 Horizon dense and MOVA models
A pull request has been submitted to the llama.cpp project to add support for the K2 Horizon dense and MOVA models. This update, proposed by user bitalov, aims to integrate these specific AI models into the llama.cpp fr…
-
Qwen4Exp optimization reduces indexer score memory in llama.cpp
A pull request has been submitted to the llama.cpp project to optimize the Qwen4Exp model. This optimization aims to reduce the memory required by the indexer score, potentially allowing the Qwen Flash Next model to use…
-
llama.cpp adds MMVQ optimization for MoE models like Qwen 35B
A pull request to the llama.cpp project introduces optimizations for Mixture of Experts (MoE) models, specifically targeting architectures like Qwen 35B A3B. This enhancement, named MMVQ, aims to improve the speed of th…
-
llamacpp adds local decision model support
The llamacpp project has integrated support for local Jevíčko-like decision models. This new capability allows for more efficient and localized processing of decision-making within large language models. The integration…
-
Qwen4Exp integrates Multi Token Prediction in llama.cpp
A pull request has been submitted to the llama.cpp project to integrate Multi Token Prediction (MTP) functionality with the Qwen4Exp model. This development, completed by am17an, allows for the use of Qwen Flash Next wi…
-
llama.cpp receives pull request for Qwen4Exp 'hc ops'
A pull request has been submitted to the llama.cpp project to add "hc ops" for Qwen4Exp. This update is expected to necessitate re-benchmarking of the Qwen Flash Next model. The contribution comes from a user named am17an.
-
llama.cpp bug causes non-deterministic results for M-RoPE embedding batches
A bug in the llama.cpp library causes incorrect results when processing embedding batches for M-RoPE models like Qwen3.5 and Qwen2.5-VL. The issue stems from a heap buffer overflow where the library reads past the alloc…
-
Maple 20B-A1B MoE architecture added to llama.cpp for low VRAM users
The llama.cpp project has integrated the Maple 20B-A1B ternary Mixture of Experts (MoE) architecture. This addition is expected to benefit users with limited VRAM, making the model more accessible for lower-end hardware…
-
llama.cpp adds AMD RDNA2 performance improvements · 1 source tracked
A pull request to the llama.cpp project, specifically for ggml-cuda, introduces performance improvements for AMD's RDNA2 architecture, including MI50 and MI60 GPUs. These enhancements are detailed in benchmarks and disc…
-
llama.cpp sees Flash Attention tuning for RDNA GPUs
A pull request has been submitted to the llama.cpp project, focusing on optimizing Flash Attention for CUDA and HIP architectures. The changes specifically target the gfx1201 hardware, with potential performance improve…
-
Spark-X2.5 LLM now runs locally with official llama.cpp support
The Spark-X2.5 large language model now officially supports local execution via llama.cpp. This enables users to run the model entirely on their CPU, with a demonstration provided for Microsoft Windows. The release incl…
-
Tencent Hy 4 model architecture support added to llama.cpp
A pull request has been submitted to the llama.cpp project to add support for Tencent's Hy 4 (hy_v4) model architecture. This integration aims to enable the use of the Hy 4 model within the llama.cpp framework, which is…
-
llama.cpp adds AVX2 support for faster prompt processing
A pull request to the llama.cpp project introduces AVX2 instruction set support to accelerate prompt processing for IQ models, particularly with large batch sizes. This optimization aims to improve the speed of local la…
-
llama.cpp adds CUDA optimizations for MoE model performance
A pull request to the llama.cpp project introduces CUDA optimizations for Mixture of Experts (MoE) models. These enhancements aim to improve performance, particularly for speculative decoding and MoE routing, by extendi…
-
llama.cpp integrates DFlash2 for improved local LLM performance
The llama.cpp project has integrated support for DFlash2, a new technique that enhances local convolution and candidate selection. This merge, identified as Pull Request #27342, was contributed by SubSir and is now part…
-
llama.cpp adds --n-cpu-ffn option for faster dense models on low VRAM
A new pull request for the llama.cpp project introduces the `--n-cpu-ffn` option, designed to improve the performance of dense models for users with limited VRAM. This feature allows a specified number of FFN sublayers …
-
llama.cpp adds support for DSpark Nanbeige4.2-3B model
A pull request has been submitted to the llama.cpp project to add support for the DSpark model, specifically the Nanbeige4.2-3B variant. This contribution, made by user zqlcode, aims to integrate the model into the llam…
-
llama.cpp PR boosts IQ model prompt processing with AVX2 optimizations · 1 source tracked
A pull request for the llama.cpp project introduces AVX2 optimizations to significantly accelerate prompt processing for IQ models, particularly at large batch sizes. Benchmarks show dramatic speed increases, with some …
-
llama.cpp adds CPU offload option for dense models
A new pull request has been submitted to the llama.cpp project, proposing the addition of a `--n-cpu-ffn` option. This feature aims to provide CPU offloading capabilities for dense models, similar to the existing `--n-c…
-
Open-source projects dotenvy and llama.cpp see significant updates
The open-source project dotenvy is being forked into a new project called dotenv-ng, with the developers citing a desire for a fresh start and improved maintenance. Separately, the llama.cpp project has released version…