PulseAugur
EN
LIVE 22:44:01

Qwen3 235B A22B model shows benchmark performance, including 70% on GPQA

The Qwen3 235B A22B model has demonstrated performance metrics across several benchmarks, including a 70% score on GPQA and 82.8% on MMLU-Pro. It achieved 11% on Humanity's Last Exam and 0% on long-context reasoning tasks. These results were independently measured and indicate a cost-effectiveness of 5.1 intelligence points per dollar. AI

IMPACT Provides specific benchmark scores for the Qwen3 235B A22B model, useful for comparing its capabilities against other LLMs.

RANK_REASON The item reports benchmark results for an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3 235B A22B model shows benchmark performance, including 70% on GPQA

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per

    📊 Qwen3 235B A22B (Reasoning) — the actual numbers GPQA: 70% MMLU-Pro: 82.8% Humanity's Last Exam: 11% Long Context Reasoning: 0% 💰 5.1 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # A…