GPQA: A Graduate-Level Google-Proof Q&A Benchmark
PulseAugur coverage of GPQA: A Graduate-Level Google-Proof Q&A Benchmark — every cluster mentioning GPQA: A Graduate-Level Google-Proof Q&A Benchmark across labs, papers, and developer communities, ranked by signal.
- instance of long-context reasoning 90%
- used by MATH500 90%
- instance of GLM-5.2 90%
- instance of MATH500 90%
- instance of MMLU-Pro 70%
- instance of Humanity's Last Exam 70%
- competes with MMLU-Pro 70%
- instance of SciCode 70%
- used by GSM8K 70%
- instance of HumanEval 70%
- instance of Artificial Intelligence In Medical Epidemiology 70%
- instance of LiveCodeBench 70%
10 day(s) with sentiment data
-
DeepSeek V4 Pro and Nova 2.0 Lite benchmarks revealed · 2 sources tracked
Independent benchmarks reveal the performance of two large language models, DeepSeek V4 Pro and Nova 2.0 Lite. DeepSeek V4 Pro achieved high scores in reasoning-focused tasks, with 92.8% on GPQA and 80.3% on Long Contex…
-
Qwen3 Omni 30B A3B Instruct shows mixed benchmark results, excelling in some areas but failing in long-context reasoning
The Qwen3 Omni 30B A3B Instruct model has demonstrated strong performance on specific benchmarks, achieving 62% on GPQA and 72.5% on MMLU-Pro. However, it showed no capability in long-context reasoning. The model operat…
-
DeepSeek-V4 Flash 0731 achieves strong GPQA and HLE scores
DeepSeek-V4 Flash 0731 has achieved a score of 90.8% on the GPQA benchmark and 38.6% on HLE. The model also demonstrated a speed of 233.8 tokens per second. Notably, it offers a high intelligence-to-cost ratio, providin…
-
Seed-OSS-36B-Instruct achieves high cost-efficiency in benchmarks
Seed-OSS-36B-Instruct has achieved a cost-efficiency of 40.3 intelligence points per dollar, according to independent measurements. The model scored 72.6% on the GPQA benchmark and 81.5% on MMLU-Pro. These results sugge…
-
Solar Pro 4 achieves 89.1% on GPQA, shows strong efficiency
Solar Pro 4 has achieved an 89.1% score on the GPQA benchmark, demonstrating strong performance in graduate-level question answering. The model also exhibited an efficiency of 59.1 tokens per second with 62.3 intelligen…
-
GLM-5.1 (Non-reasoning) achieves 83.9% on GPQA benchmark
A new benchmark evaluation shows that GLM-5.1 (Non-reasoning) achieved an 83.9% score on the GPQA benchmark. This performance was measured independently and highlights the model's efficiency, delivering 13.3 intelligenc…
-
Qwen3.5 2B shows mixed results in independent benchmarks
Independent benchmarks reveal that Qwen3.5 2B, a non-reasoning model, achieves 43.8% on the GPQA benchmark. However, its performance significantly drops on more complex tasks, scoring only 5% on HLE, 15% on Long Context…
-
Qwen3.6 and Qwen3.5 show similar inference speeds, with gains in agentic tasks
A recent benchmark comparison of Qwen3.6 and Qwen3.5 models revealed that their inference speeds on a GeForce RTX 4070 were nearly identical, contrary to initial findings that suggested a significant slowdown. This disc…
-
New research suggests RL enhances language model sampling efficiency, not new reasoning
A new research paper explores how reinforcement learning (RL) impacts language model reasoning, specifically whether it introduces new reasoning capabilities or enhances the sampling of existing ones. The study introduc…
-
BenchMIRT method reveals what LLM benchmarks truly measure · 2 sources tracked
Researchers have introduced BenchMIRT, a novel methodology designed to dissect the performance of large language models (LLMs) on benchmarks by analyzing individual prompts. This approach, inspired by Item Response Theo…
-
New PRACT-120 benchmark aims to evaluate AI chatbots holistically
A new benchmark called PRACT-120 has been proposed to evaluate AI chatbots more comprehensively than existing tests like MMLU or GPQA. The benchmark aims to assess not just the core model's capabilities but also the int…
-
LLM Benchmark Results: Molmo, Qwen3, Hermes, and Mistral Performance Revealed
Independent benchmarks reveal varying performance across several large language models. Molmo 7B-D shows low scores on GPQA and MMLU-Pro, while Qwen3 Omni 30B A3B and Hermes 4 - Llama-3.1 405B demonstrate significantly …
-
New TRACES framework enables cost-efficient early stopping for LLM reasoning
Researchers have introduced TRACES, a new framework designed to tag reasoning steps in Language Reasoning Models (LRMs) to enable adaptive and cost-efficient early stopping. This method monitors reasoning behaviors duri…
-
Kimi K2.5 achieves strong benchmark scores with competitive pricing
Kimi K2.5 has achieved notable performance on several benchmarks, including GPQA, HLE, Long Context, and SciCode. The model offers competitive pricing at 30 integer points per dollar across these evaluations. These resu…
-
LLM judges in multi-agent systems face reliability issues, new research suggests
Multiple research papers explore the limitations and potential improvements of using Large Language Models (LLMs) as judges in multi-agent systems and for evaluating agentic tool-calling. One study introduces AgentAudit…
-
New GIM benchmark evaluates LLMs on integrated cognitive tasks
Researchers have introduced the Grounded Integration Measure (GIM), a new benchmark designed to evaluate large language models (LLMs) by assessing their ability to integrate multiple cognitive operations. Unlike benchma…
-
Qwen3.5 35B A3B model achieves high GPQA score and cost-efficiency
The Qwen3.5 35B A3B model has demonstrated impressive performance, achieving an 81.9% score on the GPQA benchmark. It also offers a high intelligence-to-cost ratio, delivering 35.3 intelligence points per dollar. The mo…
-
Anthropic's Claude 3.5 Sonnet overtakes Opus in benchmarks, becomes default for builders
Anthropic has released Claude 3.5 Sonnet, a new mid-tier model that surpasses its previous flagship, Claude 3 Opus, in reasoning and coding benchmarks. This release offers a significant improvement in the cost-performan…
-
Sarvam 30B model performance metrics revealed across benchmarks
Sarvam AI has released its Sarvam 30B model, with performance metrics now available for several benchmarks. The model achieved 63.3% on GPQA, 7.5% on Humanity's Last Exam, and 19.2% on SciCode. Notably, it scored 0% on …
-
LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked
Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various…