PulseAugur
EN
LIVE 09:37:06

Grok 4.6 benchmark scores vary wildly based on reporting method

xAI's Grok 4.6 has shown vastly different performance scores on the same benchmark, depending on how it is measured. The model achieved 26% according to xAI's own model card for Terminal-Bench 3.0, but a separate analysis by Artificial Analysis reported 88.4% on a slightly different version of the benchmark. This discrepancy highlights the importance of qualifiers when reporting benchmark results. AI

IMPACT Highlights the critical need for precise reporting of AI model benchmarks to avoid misinterpretations of performance.

RANK_REASON The item discusses benchmark results for an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Grok 4.6 benchmark scores vary wildly based on reporting method

How we ranked this

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The item discusses benchmark results for an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    Grok 4.6 Scores 26% and 88% on the Same Benchmark Line xAI's own model card reports Terminal-Bench 3.0 at 26.0%. Artificial Analysis reports 88.4% on 2.1. Both

    Grok 4.6 Scores 26% and 88% on the Same Benchmark Line xAI's own model card reports Terminal-Bench 3.0 at 26.0%. Artificial Analysis reports 88.4% on 2.1. Both are the same model. The model card is unusually honest about why — the problem is everyone quoting it drops the qualifie…