PulseAugur
实时 09:49:48
English(EN) Grok 4.6 Scores 26% and 88% on the Same Benchmark Line xAI's own model card reports Terminal-Bench 3.0 at 26.0%. Artificial Analysis reports 88.4% on 2.1. Both

Grok 4.6 基准测试得分因报告方法不同而差异巨大

xAIGrok 4.6 在同一基准测试中,根据测量方式的不同,表现得分差异巨大。根据 xAI 自家的 Terminal-Bench 3.0 模型卡,该模型得分为 26%,但 Artificial Analysis 的一项独立分析报告称,在基准测试的一个略有不同的版本上得分为 88.4%。这种差异凸显了在报告基准测试结果时限定条件的重要性。 AI

影响 强调了精确报告 AI 模型基准测试结果以避免误解性能的极端重要性。

排序理由 该条目讨论了 AI 模型的基准测试结果,属于研究范畴。[lever_c_research降级:ic=1 ai=1.0]

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Grok 4.6 基准测试得分因报告方法不同而差异巨大

本文如何被排名

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目讨论了 AI 模型的基准测试结果,属于研究范畴。[lever_c_research降级:ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Grok 4.6 在同一基准测试中得分 26% 和 88% xAI 自己的模型卡报告 Terminal-Bench 3.0 为 26.0%。Artificial Analysis 报告称在 2.1 版本上得分 88.4%。两者

    Grok 4.6 Scores 26% and 88% on the Same Benchmark Line xAI's own model card reports Terminal-Bench 3.0 at 26.0%. Artificial Analysis reports 88.4% on 2.1. Both are the same model. The model card is unusually honest about why — the problem is everyone quoting it drops the qualifie…