GPQA: A Graduate-Level Google-Proof Q&A Benchmark
PulseAugur coverage of GPQA: A Graduate-Level Google-Proof Q&A Benchmark — every cluster mentioning GPQA: A Graduate-Level Google-Proof Q&A Benchmark across labs, papers, and developer communities, ranked by signal.
- instance of long-context reasoning 90%
- instance of MATH500 90%
- instance of MMLU-Pro 70%
- used by MATH500 70%
- used by GSM8K 70%
- instance of Artificial Intelligence In Medical Epidemiology 70%
- instance of GLM-5.2 70%
- instance of HumanEval 70%
- instance of Humanity's Last Exam 70%
- instance of LiveCodeBench 70%
- instance of Math-500 60%
8 day(s) with sentiment data
-
Data repetition outperforms scaling in LLM fine-tuning, study finds
A new research paper suggests that for supervised fine-tuning of reasoning language models, data repetition can be more effective than data scaling. The study found that training for more epochs on smaller datasets, eve…
-
New DRBENCHER benchmark tests AI agents' combined browsing and math skills
Researchers have introduced DRBENCHER, a new benchmark designed to evaluate AI agents' ability to combine web browsing with multi-step mathematical computations. Unlike previous benchmarks that assess these skills in is…
-
LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models
Independent benchmarks reveal varying performance across several large language models. Kimi K2 0905 achieved strong scores on GPQA and MMLU-Pro, while Qwen3 235B A22B also performed well on these metrics but struggled …
-
Korean AI models face scrutiny over 'borrowed' tech and narrow benchmarks · 1 source tracked
A viral debate on China's Zhihu platform scrutinized a wave of Korean AI models that claimed to outperform DeepSeek. The critique focused on several points: models allegedly building upon open-weight models from other l…
-
Divergence Decoding fuses LLM capabilities without retraining
Researchers have introduced Divergence Decoding, a novel training-free framework designed to fuse the capabilities of specialized scientific language models with generalist models. This method uses Jensen-Shannon diverg…
-
AI agent performance hinges on outcomes, not just benchmarks
Benchmarks like MMLU and GPQA do not accurately reflect real-world performance for AI agents, according to Aysan Isayo. While benchmarks focus on the correctness of answers, the actual utility of an agent depends more o…
-
AI benchmark contamination traced to pipeline error; scores withdrawn
A pipeline error led to the contamination of four AI benchmarks, including GPQA. The issue stemmed from the GPQA evaluation set being published with a 'train' label on Hugging Face, causing a pipeline to incorrectly sel…
-
LLM benchmark results reveal performance across multiple models · 9 sources tracked
A recent independent benchmark evaluation has revealed performance metrics for several large language models, including Kimi K2, Sarvam Maya, NVIDIA Nemotron 3 Super 120B, DeepSeek V3.2, Falcon H1R-7B, GLM-5.2, GLM-5.1,…
-
Soofi consortium withdraws GPQA benchmark scores due to data contamination
The Soofi consortium in Germany has retracted its benchmark scores due to the discovery that its training data included paraphrased questions from the GPQA evaluation set. This contamination means the previously reporte…
-
DBRX Instruct and Mistral Medium 3 benchmark results revealed
Independent benchmarks reveal performance metrics for two large language models. DBRX Instruct achieved scores of 33.1% on GPQA, 39.7% on MMLU-Pro, 6.6% on Humanity's Last Exam, and 9.3% on LiveCodeBench. Mistral Medium…
-
New Step-Tagging framework enhances control over Language Reasoning Models
Researchers have introduced a new framework called Step-Tagging to better control the generation process of Language Reasoning Models (LRMs). This framework uses a lightweight sentence classifier to annotate reasoning s…
-
Open-source LLMs show strong benchmark performance across multiple metrics · 4 sources tracked
Several open-source AI models have demonstrated strong performance on various benchmarks, according to independent measurements. Mi:dm K 2.5 Pro achieved 70.1% on GPQA and 80.9% on MMLU-Pro, while MiMo-V2-Flash showed 8…
-
New 'Representation-as-a-Judge' method uses small models for evaluation
Researchers have proposed a new evaluation method for language models called Representation-as-a-Judge, which utilizes the internal representations of smaller models rather than their generative output. This approach is…
-
Qwen 3.6 quantizations show agentic performance drop, knowledge recall stable
A university HPC cluster has benchmarked Qwen 3.6 quantizations, revealing that lower-precision versions significantly degrade agentic performance as measured by Terminal-Bench 2. While knowledge recall, assessed by GPQ…
-
New method penalizes redundancy to make LLM reasoning more efficient
Researchers have developed a novel method to reduce "overthinking" in large reasoning models (LRMs) by penalizing both internal and external redundancy in their Chain-of-Thought (CoT) traces. This dual-penalty reinforce…
-
New 'LearnStop' method optimizes reasoning model stopping points
Researchers have developed a new method called LearnStop to optimize when reasoning language models should stop processing an instance. This technique analyzes multiple features like answer confidence, entropy, and stab…
-
New research challenges on-policy self-distillation for LLMs, proposing refined methods · 10 sources tracked
Recent research papers explore the limitations and potential improvements of on-policy self-distillation (OPSD) for training large language models (LLMs). Studies indicate that standard OPSD can lead to rote memorizatio…
-
Zhipu AI's GLM-5.2 model deployed on serverless GPUs
Zhipu AI has released GLM-5.2, a 700B Mixture-of-Experts (MoE) model that excels in complex reasoning and software engineering tasks, reportedly matching or surpassing proprietary models like Claude 3.5 Sonnet and GPT-4…
-
New RAD method controls MoE language model reasoning without text analysis
Researchers have developed a new method called RAD (Routing Agreement Decoding) for controlling reasoning in sparse Mixture-of-Experts (MoE) language models. This technique leverages the internal routing states of MoE m…
-
New SIGMA framework boosts AI mathematical reasoning with multi-agent knowledge integration
Researchers have developed SIGMA, a novel framework designed to improve mathematical reasoning in AI agents. SIGMA employs a multi-agent system where specialized agents independently reason, conduct targeted searches, a…