The Qwen3 235B A22B model has demonstrated performance metrics across several benchmarks, including a 70% score on GPQA and 82.8% on MMLU-Pro. It achieved 11% on Humanity's Last Exam and 0% on long-context reasoning tasks. These results were independently measured and indicate a cost-effectiveness of 5.1 intelligence points per dollar. AI
IMPACT Provides specific benchmark scores for the Qwen3 235B A22B model, useful for comparing its capabilities against other LLMs.
RANK_REASON The item reports benchmark results for an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- long-context reasoning
- MMLU-Pro
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →