Large language models (LLMs) set to temperature 0, which typically makes them greedy and deterministic, can still produce different outputs for the same prompt. This variability arises not from the model's creativity but from the underlying hardware and software execution. Specifically, floating-point arithmetic's non-associative nature and variations in GPU kernel reduction orders, influenced by batch sizes and other concurrent requests, lead to subtle shifts in logits. These minor numerical differences can be amplified during autoregressive decoding, resulting in altered output, a phenomenon known as a lack of batch invariance. AI
IMPACT Highlights the challenges of achieving bitwise reproducibility in LLM inference, impacting testing and deployment.
RANK_REASON Technical explanation of non-determinism in LLM inference. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →