PulseAugur
实时 01:46:14
English(EN) Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Termi

SemiAnalysis 表示,公开的 AI 基准测试会因模型对其进行优化而迅速失去价值

SemiAnalysis 认为,像 TB 4.0 这样的公开 AI 基准测试会因模型对其进行优化而迅速过时,从而降低了它们评估真正泛化能力的有效性。他们以 Gemini 3.8 FlashMuse Spark 1.3 为例,这些模型在 Terminal Bench 2.1 等旧基准测试上表现良好,但在 TB 4.0 等新基准测试上表现不佳,这表明私有的高质量基准测试是评估模型能力的更可靠的解决方案。分析还指出,像 Datacurve 这样的公司通过创建模仿这些基准测试的任务来为大型 AI 实验室牟利。 AI

影响 建议转向私有基准测试来评估 AI 模型,这可能会影响衡量和比较 AI 能力的方式。

排序理由 该集群包含 SemiAnalysis 的观点文章,讨论了公开 AI 基准测试的局限性以及对私有基准测试的需求。

在 X — SemiAnalysis 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

SemiAnalysis 表示,公开的 AI 基准测试会因模型对其进行优化而迅速失去价值

本文如何被排名

Signal score
3 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
该集群包含 SemiAnalysis 的观点文章,讨论了公开 AI 基准测试的局限性以及对私有基准测试的需求。
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [5]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    解决方案是更高质量的私有基准测试。如果您有一个出色的私有基准测试,请联系 @maxkan。我们很乐意与您交流!(5/5)

    The solution is more high quality private benchmarks. If you have a great private benchmark, reach out to @maxkan. We'd love to chat! (5/5)

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Gemini 3.8 Flash on DeepSWE 是另一个好例子。显然,Datacurve 通过销售 Google DeepSWE 形式的任务赚了很多钱。(4/5) https://t.co/S6lf25GofC

    Gemini 3.8 Flash on DeepSWE is another good example. Clearly, Datacurve made a ton of money selling Google DeepSWE-shaped tasks. (4/5) https://t.co/S6lf25GofC

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    最终,这就是所有优秀公开基准的命运。TB 4.0 也不例外。它现在之所以还有用,仅仅是因为它是在两周前发布的。因为所有 t

    Ultimately, this is the fate of all good public benchmarks. TB 4.0 is no exception. It’s only useful signal now because it was released 2 weeks ago. Since all the tasks are similarly public, it won't be long until it's hillclimbed by all the aspiring “frontier” labs. (3/5)

  4. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    这怎么可能?Terminal Bench 2.1 中的所有任务都是完全公开的。尽管 Meta 和 Google 绝不会直接在这些任务上进行训练,但它们绝对会

    How is this possible? All of the tasks in Terminal Bench 2.1 are fully public. Though Meta and Google would never train on the tasks directly, they absolutely will buy data from RL env startups that’s designed to mimic TB 2.1 tasks as closely as possible. The net effect is the

  5. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    Gemini 3.8 Flash 和 Muse Spark 1.3 是我们迄今为止见过的最明显经过跑分优化的模型。尽管它们在 Termi 上可与 GPT-6 和 Fable 5.1 相媲美。

    Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse. (1/5)🧵 https://t.co/K2ccQjWm11