Hugging Face has launched a new interactive map consolidating its 48 official benchmarks, providing a comprehensive overview of AI model performance across various domains. The map reveals that while agent-related benchmarks are most numerous, science and knowledge benchmarks attract the most submissions due to practical considerations. Participation is uneven, with a few key benchmarks dominating leaderboard entries, and a significant portion of submissions originating from China. AI
IMPACT Provides a consolidated view of AI model performance, aiding researchers and developers in understanding benchmark landscape and participation trends.
RANK_REASON Launch of a new interactive map/tool by Hugging Face to visualize existing benchmarks.
- AIME 2026
- GPQA Diamond
- Hugging Face
- Hugging Face Official Benchmarks: Map and Participation
- MMLU-Pro
- Open ASR Leaderboard
- SWE-bench
- SWE-bench Verified
- Terminal-Bench
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →