📊 Qwen3 Max — the actual numbers GPQA: 76.4% MMLU-Pro: 84.1% Humanity's Last Exam: 11.1% Long Context Reasoning: 46.7% 💰 10 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI
📊 Kimi K2 — the actual numbers GPQA: 76.6% MMLU-Pro: 82.4% Humanity's Last Exam: 7% Long Context Reasoning: 51% 💰 19.4 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI
📊 Nova 2.0 Lite (medium) — the actual numbers GPQA: 76.8% MMLU-Pro: 81.3% Humanity's Last Exam: 8.6% Long Context Reasoning: 58.3% ⚡ 217.9 tokens/sec 💰 22.4 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchm…
📊 Sarvam M (Reasoning) — the actual numbers GPQA: 41.6% MMLU-Pro: 69.6% Humanity's Last Exam: 3.3% Long Context Reasoning: 0% Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI
📊 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) — the actual numbers GPQA: 80% Humanity's Last Exam: 19.2% Long Context Reasoning: 60% SciCode: 36% ⚡ 157.5 tokens/sec 💰 66.7 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/lead…
📊 DeepSeek V3.2 (Reasoning) — the actual numbers GPQA: 84% MMLU-Pro: 86.2% Humanity's Last Exam: 22.2% Long Context Reasoning: 65% 💰 101.6 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # …
📊 Falcon-H1R-7B — the actual numbers GPQA: 66.1% MMLU-Pro: 72.5% Humanity's Last Exam: 10.8% Long Context Reasoning: 8.7% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI
📊 GLM-5.2 (max) — the actual numbers GPQA: 89.5% Humanity's Last Exam: 40.1% Long Context Reasoning: 71.3% SciCode: 50.5% ⚡ 156.7 tokens/sec 💰 23.8 intelligence points per dollar Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benc…
📊 GLM-5.1 (Reasoning) — the actual numbers GPQA: 86.8% Humanity's Last Exam: 28% Long Context Reasoning: 62.3% SciCode: 43.8% 💰 18.8 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI
📊 Solar Open 100B (Reasoning) — the actual numbers GPQA: 65.7% Humanity's Last Exam: 9.2% Long Context Reasoning: 36% SciCode: 26.9% Measured independently, not self-reported → https:// opensourceai.tech/leaderboard. html # LLM # Benchmarks # OpenSource # AI