PulseAugur
中
实时 14:11:03
English(EN) Hugging Face official benchmarks: the complete list (48) and how their leaderboards work

Hugging Face 排行榜:基准测试如何被填充和查看

Hugging Face 维护着 48 个官方基准测试的排行榜,这些排行榜通过模型仓库中的 .eval_results YAML 文件由模型作者提交结果来填充。这些排行榜由 Hugging Face 自动组装,并非由基准测试所有者直接上传。大约 30% 的条目在默认视图中不可见,需要进行特定过滤。 “Agents and terminal”类别包含的基准测试最多,而 “Science and knowledge”类别包含的条目最多,其中 MMLU-Pro、GPQA Diamond 和 Humanity's Last Exam 是条目最多的基准测试之一。 AI

影响 提供了关于如何在各种基准测试中跟踪和比较 AI 模型性能的见解。

排序理由 文章详细介绍了 Hugging Face 基准测试排行榜的机制和数据结构,包括分数如何提交和显示。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

Hugging Face 排行榜:基准测试如何被填充和查看

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
文章详细介绍了 Hugging Face 基准测试排行榜的机制和数据结构,包括分数如何提交和显示。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
3 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [3]

  1. X — Aravind Srinivas (Perplexity) TIER_1 English(EN) · AravSrinivas ·

    Perplexity 在 Hugging Face 决策模型指数(决策模型基准)中胜出。我们的模型是开放权重且人人可用的!

    Perplexity wins on Hugging Face Decision Index (benchmark for decision models). Our model is open-weights and accessible to all!

  2. dev.to — LLM tag TIER_1 English(EN) · Ward Ed ·

    Hugging Face 官方基准排行榜的实际运作方式:.eval_results YAML、base_model 过滤器以及默认视图隐藏的 30%

    <p><strong>TL;DR:</strong> Hugging Face currently tags 48 datasets as official benchmarks (<code>benchmark:official</code>). Their leaderboards are not uploaded by the benchmark owners: they are assembled automatically from small <code>.eval_results/*.yaml</code> files that model…

  3. dev.to — LLM tag TIER_1 English(EN) · AI OpenFree ·

    Hugging Face 官方基准测试:完整列表(48个)及其排行榜工作原理

    <p><strong>Short answer:</strong> Hugging Face currently marks <strong>48 datasets as official benchmarks</strong> (October 4, 2026). Each has a leaderboard on its dataset page, built from <code>.eval_results</code> files that model repositories publish. Together they hold about …