PulseAugur
EN
LIVE 10:33:08

LLM benchmark results reveal performance across multiple models · 9 sources tracked

A recent independent benchmark evaluation has revealed performance metrics for several large language models, including Kimi K2, Sarvam Maya, NVIDIA Nemotron 3 Super 120B, DeepSeek V3.2, Falcon H1R-7B, GLM-5.2, GLM-5.1, Solar Open 100B, and GLM-4.7-Flash. The benchmarks cover areas such as reasoning, MMLU-Pro, Humanity's Last Exam, and long-context reasoning, with results varying significantly across models. Notably, GLM-5.2 and DeepSeek V3.2 show strong performance in reasoning and MMLU-Pro, while Kimi K2 and Nemotron 3 Super 120B also demonstrate competitive scores. AI

IMPACT Provides comparative performance data for various LLMs across key benchmarks, aiding developers in model selection.

RANK_REASON The cluster reports benchmark results for multiple LLMs, which falls under research.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 12 sources. How we write summaries →

LLM benchmark results reveal performance across multiple models · 9 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
The cluster reports benchmark results for multiple LLMs, which falls under research.
Source corroboration
12 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
45 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [12]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 Max — the actual numbers GPQA: 76.4% MMLU-Pro: 84.1% Humanity's Last Exam: 11.1% Long Context Reasoning: 46.7% 💰 10 intelligence points per dollar Measu

    📊 Qwen3 Max — the actual numbers GPQA: 76.4% MMLU-Pro: 84.1% Humanity's Last Exam: 11.1% Long Context Reasoning: 46.7% 💰 10 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Kimi K2 — the actual numbers GPQA: 76.6% MMLU-Pro: 82.4% Humanity's Last Exam: 7% Long Context Reasoning: 51% 💰 19.4 intelligence points per dollar Measured i

    📊 Kimi K2 — the actual numbers GPQA: 76.6% MMLU-Pro: 82.4% Humanity's Last Exam: 7% Long Context Reasoning: 51% 💰 19.4 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Nova 2.0 Lite (medium) — the actual numbers GPQA: 76.8% MMLU-Pro: 81.3% Humanity's Last Exam: 8.6% Long Context Reasoning: 58.3% ⚡ 217.9 tokens/sec 💰 22.4 int

    📊 Nova 2.0 Lite (medium) — the actual numbers GPQA: 76.8% MMLU-Pro: 81.3% Humanity's Last Exam: 8.6% Long Context Reasoning: 58.3% ⚡ 217.9 tokens/sec 💰 22.4 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchm…

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 5.1% Long Context Reasoning: 4% Measured independently, not self-reporte

    📊 Granite 4.0 1B — the actual numbers GPQA: 28.1% MMLU-Pro: 32.5% Humanity's Last Exam: 5.1% Long Context Reasoning: 4% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Sarvam M (Reasoning) — the actual numbers GPQA: 41.6% MMLU-Pro: 69.6% Humanity's Last Exam: 3.3% Long Context Reasoning: 0% Measured independently, not self-r

    📊 Sarvam M (Reasoning) — the actual numbers GPQA: 41.6% MMLU-Pro: 69.6% Humanity's Last Exam: 3.3% Long Context Reasoning: 0% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  6. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) — the actual numbers GPQA: 80% Humanity's Last Exam: 19.2% Long Context Reasoning: 60% SciCode: 36% ⚡ 157.5 toke

    📊 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) — the actual numbers GPQA: 80% Humanity's Last Exam: 19.2% Long Context Reasoning: 60% SciCode: 36% ⚡ 157.5 tokens/sec 💰 66.7 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/lead…

  7. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 DeepSeek V3.2 (Reasoning) — the actual numbers GPQA: 84% MMLU-Pro: 86.2% Humanity's Last Exam: 22.2% Long Context Reasoning: 65% 💰 101.6 intelligence points p

    📊 DeepSeek V3.2 (Reasoning) — the actual numbers GPQA: 84% MMLU-Pro: 86.2% Humanity's Last Exam: 22.2% Long Context Reasoning: 65% 💰 101.6 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # …

  8. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Falcon-H1R-7B — the actual numbers GPQA: 66.1% MMLU-Pro: 72.5% Humanity's Last Exam: 10.8% Long Context Reasoning: 8.7% Measured independently, not self-repor

    📊 Falcon-H1R-7B — the actual numbers GPQA: 66.1% MMLU-Pro: 72.5% Humanity's Last Exam: 10.8% Long Context Reasoning: 8.7% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  9. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 GLM-5.2 (max) — the actual numbers GPQA: 89.5% Humanity's Last Exam: 40.1% Long Context Reasoning: 71.3% SciCode: 50.5% ⚡ 156.7 tokens/sec 💰 23.8 intelligence

    📊 GLM-5.2 (max) — the actual numbers GPQA: 89.5% Humanity's Last Exam: 40.1% Long Context Reasoning: 71.3% SciCode: 50.5% ⚡ 156.7 tokens/sec 💰 23.8 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benc…

  10. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 GLM-5.1 (Reasoning) — the actual numbers GPQA: 86.8% Humanity's Last Exam: 28% Long Context Reasoning: 62.3% SciCode: 43.8% 💰 18.8 intelligence points per dol

    📊 GLM-5.1 (Reasoning) — the actual numbers GPQA: 86.8% Humanity's Last Exam: 28% Long Context Reasoning: 62.3% SciCode: 43.8% 💰 18.8 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI

  11. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Solar Open 100B (Reasoning) — the actual numbers GPQA: 65.7% Humanity's Last Exam: 9.2% Long Context Reasoning: 36% SciCode: 26.9% Measured independently, not

    📊 Solar Open 100B (Reasoning) — the actual numbers GPQA: 65.7% Humanity's Last Exam: 9.2% Long Context Reasoning: 36% SciCode: 26.9% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  12. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 GLM-4.7-Flash (Non-reasoning) — the actual numbers GPQA: 45.2% Humanity's Last Exam: 4.9% Long Context Reasoning: 14.7% SciCode: 25.5% ⚡ 179.6 tokens/sec 💰 10

    📊 GLM-4.7-Flash (Non-reasoning) — the actual numbers GPQA: 45.2% Humanity's Last Exam: 4.9% Long Context Reasoning: 14.7% SciCode: 25.5% ⚡ 179.6 tokens/sec 💰 101.3 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. h…