A user experimenting with OpenAI's models discovered that even with temperature set to 0.0, the output for a safety-related scoring task was inconsistent. Running the same prompt and identical images multiple times resulted in the model's classification flipping between 'Low' and 'Medium' risk categories. This suggests that factors beyond temperature, such as internal routing or floating-point arithmetic, can introduce non-determinism in model outputs, especially when scores fall near a decision threshold. AI
IMPACT Highlights potential issues with model reproducibility for safety-critical applications, even with strict parameter settings.
RANK_REASON User-generated observation about model behavior, not a direct release or research paper.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →