PulseAugur
EN
LIVE 04:04:40

User benchmarks show mainline llama.cpp outperforms ik_llama.cpp in multi-GPU setups

A user on r/LocalLLaMA is questioning the reported speed advantage of the ik_llama.cpp fork over the mainline llama.cpp project, particularly in hybrid multi-GPU setups. Benchmarks on the user's system, which includes an Intel Core i9-10920X and four NVIDIA GPUs, show mainline significantly outperforming ik_llama.cpp in both prompt evaluation and token generation speeds. The user suspects ik_llama.cpp might be optimized for CPU-only or single-GPU inference, and its custom threadpool may not scale well with pipelined layer splits across heterogeneous GPUs, unlike mainline's CUDA graph caching. AI

IMPACT Performance differences in local LLM inference frameworks can impact user experience and hardware optimization choices.

RANK_REASON User-generated technical discussion and benchmark comparison of two software forks.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

User benchmarks show mainline llama.cpp outperforms ik_llama.cpp in multi-GPU setups

How we ranked this

Signal score
2 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
User-generated technical discussion and benchmark comparison of two software forks.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/vulcan4d ·

    Why is ik_llama.cpp said to be faster than Mainline? On my hybrid multi-GPU rig, Mainline easily beats it

    <!-- SC_OFF --><div class="md"><p>I constantly see recommendations saying that ik_llama.cpp (ikawrakow's fork) is the undisputed king of hybrid CPU/GPU offloading and MoE performance. However, every time I benchmark it against mainline ggml-org, mainline consistently beats it by …