PulseAugur
EN
LIVE 04:10:27

Llama-server prompt cache bug causes incorrect LoRA scale inference

A bug in the `llama-server`'s RAM prompt cache can lead to incorrect text generation when using LoRA adapters at different scales. The cache stores computed KV states but does not record the specific LoRA adapter and scale used, causing the server to potentially reuse outdated KV states. This issue was observed in multiple versions of `llama.cpp`, with a 100% failure rate in testing when a prompt was first processed with one LoRA scale and then re-processed with a different scale after another request occupied the processing slot. Disabling the RAM cache or setting `cache_prompt: false` for each request resolves the problem. AI

IMPACT This bug could lead to inconsistent or nonsensical outputs for users employing LoRA adapters with `llama-server`, potentially impacting applications that rely on fine-tuned model behavior.

RANK_REASON Bug report in a specific feature of an open-source LLM serving tool.

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Llama-server prompt cache bug causes incorrect LoRA scale inference

How we ranked this

Signal score
27 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Bug report in a specific feature of an open-source LLM serving tool.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · The Homelab Postmortem ·

    llama-server's prompt cache reuses KV computed under a different LoRA scale

    <p><strong>TL;DR</strong>: <code>llama-server</code>'s RAM prompt cache (<code>--cache-ram</code>, on by default) saves a slot's KV when another request takes the slot, and restores it when the same prompt comes back. The saved entry doesn't record which LoRA adapters and scales …