A bug in llama-server's sleep mode can cause requests to be lost or lead to server crashes. When the server enters its sleep state, a race condition can occur if a request arrives just before or during this transition. Requests arriving too close to the sleep initiation may not be processed, leading to a hang until a subsequent request wakes the server, or in some cases, a SIGSEGV crash within the tokenizer when it attempts to access vocabulary that has been unloaded. This issue has been observed with Gemma 3:1B models on llama.cpp version b11368 and has been reported upstream. AI
IMPACT This bug could impact the reliability of self-hosted LLM deployments using llama-server, potentially leading to dropped user requests or service interruptions.
RANK_REASON Bug report for a specific feature of an open-source LLM serving tool.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →