A bug in the llama-server software causes speculative decoding to incorrectly report log probabilities as zero for generated tokens. This issue affects various models and configurations, including Gemma 3B, when speculative decoding is enabled. While the generated text appears normal, the log probability data is inaccurate, potentially impacting downstream applications that rely on these values. The problem is not easily detectable as zero log probabilities can occur naturally, and the server does not indicate that these are placeholder values. AI
IMPACT This bug could lead to inaccurate performance metrics and affect applications relying on precise log probability data from LLM outputs.
RANK_REASON Bug report in an open-source LLM serving tool.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →