Hugging Face has officially recognized 48 datasets as benchmarks as of October 4, 2026, each with its own leaderboard. These leaderboards aggregate results from approximately 410 models submitted by 95 organizations. The 'Agents and terminal' category hosts the most benchmarks with 17, while 'Science and knowledge' has the most entries. Prominent contributors to these leaderboards include Qwen, DeepSeek, Moonshot AI, and OpenAI. AI
IMPACT Provides a consolidated view of AI model performance across various benchmarks, aiding in comparative analysis.
RANK_REASON Article details a list of official benchmarks and participation metrics on a platform, not a new model release or research paper. [lever_c_demoted from research: ic=1 ai=1.0]
- DeepSeek
- GPQA Diamond
- Hugging Face
- Humanity's Last Exam
- MMLU-Pro
- Moonshot AI
- Nvidia
- OpenAI
- Open ASR Leaderboard
- Ornithogalum
- Qwen
- SWE-bench Verified
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →