The Soofi consortium in Germany has retracted its benchmark scores due to the discovery that its training data included paraphrased questions from the GPQA evaluation set. This contamination means the previously reported 11.1-point gain on the GPQA-Diamond benchmark is no longer a valid measure of model capability. A researcher identified the issue, and the Soofi team confirmed the data contamination within a week. AI
IMPACT Data contamination in training sets can invalidate benchmark results, highlighting the need for rigorous evaluation and data hygiene in AI research.
RANK_REASON Withdrawal of benchmark scores due to data contamination in a research evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →