A pipeline error led to the contamination of four AI benchmarks, including GPQA. The issue stemmed from the GPQA evaluation set being published with a 'train' label on Hugging Face, causing a pipeline to incorrectly select it based on name rather than its intended meaning. This oversight has led to the withdrawal of scores from these benchmarks, and future datasets will undergo allowlist validation and n-gram screening to prevent recurrence. AI
IMPACT Errors in benchmark datasets can skew performance evaluations, potentially misdirecting research and development efforts.
RANK_REASON The item discusses an issue with AI benchmark datasets and evaluation pipelines, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →