A new audit of the Massive Multitask Language Understanding (MMLU) benchmark reveals that its aggregate score primarily measures factual retrieval rather than reasoning capabilities. Researchers used Item Response Theory to analyze 14,042 MMLU test items across 1,000 language models, finding that the benchmark conflates distinct constructs and that difficulty varies significantly between STEM and non-STEM partitions. This imbalance inadvertently favors models optimized for retrieval, potentially misguiding the selection of models for reasoning-intensive tasks. AI
IMPACT Highlights limitations in current LLM evaluation, suggesting a need for more nuanced benchmarks that accurately assess reasoning abilities.
RANK_REASON Academic paper analyzing a benchmark's methodology and findings. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Hugging Face
- item response theory
- Massive Multitask Language Understanding
- Science Technology Engineering Mathematics
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →