A bug in the `llama-server`'s RAM prompt cache can lead to incorrect text generation when using LoRA adapters at different scales. The cache stores computed KV states but does not record the specific LoRA adapter and scale used, causing the server to potentially reuse outdated KV states. This issue was observed in multiple versions of `llama.cpp`, with a 100% failure rate in testing when a prompt was first processed with one LoRA scale and then re-processed with a different scale after another request occupied the processing slot. Disabling the RAM cache or setting `cache_prompt: false` for each request resolves the problem. AI
IMPACT This bug could lead to inconsistent or nonsensical outputs for users employing LoRA adapters with `llama-server`, potentially impacting applications that rely on fine-tuned model behavior.
RANK_REASON Bug report in a specific feature of an open-source LLM serving tool.
- Debian 13 LXC
- ggml-org/llama.cpp
- Hugging Face
- llama-server
- LoRA
- moe_shakespeare15M.gguf
- stories15M_MOE-F16.gguf
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →