Independent benchmarks reveal varying performance across several large language models. Kimi K2 0905 achieved strong scores on GPQA and MMLU-Pro, while Qwen3 235B A22B also performed well on these metrics but struggled with long-context reasoning. Exaone 4.0 showed lower scores across the board, particularly in reasoning tasks. Qwen3 VL 4B demonstrated a notable gap between its general knowledge capabilities and its performance on difficult reasoning challenges. AI
IMPACT These benchmarks highlight performance differences in reasoning and context handling, guiding future model development and selection.
RANK_REASON The cluster reports benchmark results for multiple LLMs, which falls under research.
Read on Mastodon — fosstodon.org →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- long-context reasoning
- MMLU-Pro
- Exaone 4.0
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →