PulseAugur
EN
LIVE 11:38:50

Soofi consortium withdraws GPQA benchmark scores due to data contamination

The Soofi consortium in Germany has retracted its benchmark scores due to the discovery that its training data included paraphrased questions from the GPQA evaluation set. This contamination means the previously reported 11.1-point gain on the GPQA-Diamond benchmark is no longer a valid measure of model capability. A researcher identified the issue, and the Soofi team confirmed the data contamination within a week. AI

IMPACT Data contamination in training sets can invalidate benchmark results, highlighting the need for rigorous evaluation and data hygiene in AI research.

RANK_REASON Withdrawal of benchmark scores due to data contamination in a research evaluation. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Soofi consortium withdraws GPQA benchmark scores due to data contamination

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions from the GPQA evaluation set. The 11.1-

    Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions from the GPQA evaluation set. The 11.1-point gain on GPQA-Diamond can no longer stand as evidence of model capability. A researcher spotted the contamination; …