An experiment involving rewriting an exam question fifty times revealed an error in the original answer key. The question, which asked about matching lid sizes to bottle orders, was incorrectly keyed to assume a specific lid size based on the bottle order. The corrected answer key now requires the system to ask for clarification when the lid size is ambiguous, a standard practice across other questions in the exam. This correction altered the scoring of previous model runs, showing that the previously favored model made a risky guess while the less favored model correctly asked for clarification. AI
IMPACT Highlights the importance of accurate data and answer keys in evaluating LLM performance and understanding their decision-making processes.
RANK_REASON The item discusses an experiment with LLMs on an exam question, focusing on the process and a correction to the answer key, rather than a new model release or significant industry event.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →