PulseAugur
EN
LIVE 08:49:44

Open-source LLMs show strong benchmark performance across multiple metrics · 4 sources tracked

Several open-source AI models have demonstrated strong performance on various benchmarks, according to independent measurements. Mi:dm K 2.5 Pro achieved 70.1% on GPQA and 80.9% on MMLU-Pro, while MiMo-V2-Flash showed 83.5% on GPQA and 20% on Humanity's Last Exam. Qwen3.5 122B A10B reached 85.7% on GPQA and 23.4% on Humanity's Last Exam, and Apriel-v1.6-15B-Thinker scored 73.3% on GPQA and 79% on MMLU-Pro. These results highlight the rapid progress in open-source LLM capabilities across different evaluation metrics. AI

IMPACT Demonstrates significant advancements in open-source LLM capabilities across various benchmarks, potentially accelerating adoption.

RANK_REASON Cluster reports benchmark results for multiple open-source LLMs, sourced from independent measurements.

Read on Mastodon — sigmoid.social →

AI-generated summary · Google Gemini · from 7 sources. How we write summaries →

Open-source LLMs show strong benchmark performance across multiple metrics · 4 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Cluster reports benchmark results for multiple open-source LLMs, sourced from independent measurements.
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
45 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+3 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

Full methodology in our editorial standards.

COVERAGE [7]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Mi:dm K 2.5 Pro — the actual numbers GPQA: 70.1% MMLU-Pro: 80.9% Humanity's Last Exam: 7.7% Long Context Reasoning: 9% Measured independently, not self-report

    📊 Mi:dm K 2.5 Pro — the actual numbers GPQA: 70.1% MMLU-Pro: 80.9% Humanity's Last Exam: 7.7% Long Context Reasoning: 9% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 K-EXAONE (Reasoning) — the actual numbers GPQA: 78.3% MMLU-Pro: 83.8% Humanity's Last Exam: 13.1% Long Context Reasoning: 55.7% Measured independently, not se

    📊 K-EXAONE (Reasoning) — the actual numbers GPQA: 78.3% MMLU-Pro: 83.8% Humanity's Last Exam: 13.1% Long Context Reasoning: 55.7% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  3. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 VL 8B (Reasoning) — the actual numbers GPQA: 57.9% MMLU-Pro: 74.9% Humanity's Last Exam: 3.3% Long Context Reasoning: 31% ⚡ 128.9 tokens/sec 💰 16.1 inte

    📊 Qwen3 VL 8B (Reasoning) — the actual numbers GPQA: 57.9% MMLU-Pro: 74.9% Humanity's Last Exam: 3.3% Long Context Reasoning: 31% ⚡ 128.9 tokens/sec 💰 16.1 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LL…

  4. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Llama 3.3 Nemotron Super 49B v1 (Reasoning) — the actual numbers GPQA: 64.3% MMLU-Pro: 78.5% Humanity's Last Exam: 6.5% Long Context Reasoning: 17% Measured i

    📊 Llama 3.3 Nemotron Super 49B v1 (Reasoning) — the actual numbers GPQA: 64.3% MMLU-Pro: 78.5% Humanity's Last Exam: 6.5% Long Context Reasoning: 17% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI

  5. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 MiMo-V2-Flash (Feb 2026) — the actual numbers GPQA: 83.5% Humanity's Last Exam: 20% Long Context Reasoning: 64.3% SciCode: 38.3% 💰 221.3 intelligence points p

    📊 MiMo-V2-Flash (Feb 2026) — the actual numbers GPQA: 83.5% Humanity's Last Exam: 20% Long Context Reasoning: 64.3% SciCode: 38.3% 💰 221.3 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # …

  6. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Qwen3.5 122B A10B (Reasoning) — the actual numbers GPQA: 85.7% Humanity's Last Exam: 23.4% Long Context Reasoning: 66.7% SciCode: 42% ⚡ 147.1 tokens/sec 💰 29.

    📊 Qwen3.5 122B A10B (Reasoning) — the actual numbers GPQA: 85.7% Humanity's Last Exam: 23.4% Long Context Reasoning: 66.7% SciCode: 42% ⚡ 147.1 tokens/sec 💰 29.4 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. htm…

  7. Mastodon — mastodon.social TIER_1 English(EN) · opensourceaitech ·

    📊 Apriel-v1.6-15B-Thinker — the actual numbers GPQA: 73.3% MMLU-Pro: 79% Humanity's Last Exam: 9.8% Long Context Reasoning: 50.3% Measured independently, not se

    📊 Apriel-v1.6-15B-Thinker — the actual numbers GPQA: 73.3% MMLU-Pro: 79% Humanity's Last Exam: 9.8% Long Context Reasoning: 50.3% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI