Humanity's Last Exam
PulseAugur coverage of Humanity's Last Exam — every cluster mentioning Humanity's Last Exam across labs, papers, and developer communities, ranked by signal.
- instance of MMLU-Pro 90%
- instance of Deep Research 90%
- instance of Claude Opus-5 90%
- instance of long-context reasoning 70%
- instance of SciCode 70%
- competes with Claude Fable-5 70%
- instance of GPQA Diamond 70%
- competes with GPT 5.6 "Sol" 70%
- instance of DeepSearchQA 70%
- instance of BrowseComp+ 70%
- instance of ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems 70%
- competes with Claude Opus-5 70%
9 day(s) with sentiment data
-
AI oversight monitors improve with answer access, but reasoning verification remains weak
A new research paper titled "The Answer Is Not the Argument" explores the effectiveness of chain-of-thought monitoring for AI oversight. The study found that providing AI monitors with a trusted reference answer signifi…
-
Anthropic launches Claude Fable 5.1 with improved coding, lower costs · 10 sources tracked
Anthropic has released its latest AI models, Claude Fable 5.1 and Mythos 5.1, which offer improved performance in coding and scientific research. Fable 5.1, the generally available version, shows significant gains on be…
-
Anthropic's Claude Fable 5.1 released with benchmark gains, cost debates · 10 sources tracked
Anthropic has released Claude Fable 5.1, a new iteration of its AI model, which is now available on platforms like Cursor and Perplexity. This release boasts significant improvements in benchmarks, particularly in agent…
-
LLM Benchmark Results: Molmo, Qwen3, Hermes, and Mistral Performance Revealed
Independent benchmarks reveal varying performance across several large language models. Molmo 7B-D shows low scores on GPQA and MMLU-Pro, while Qwen3 Omni 30B A3B and Hermes 4 - Llama-3.1 405B demonstrate significantly …
-
LLM judges in multi-agent systems face reliability issues, new research suggests
Multiple research papers explore the limitations and potential improvements of using Large Language Models (LLMs) as judges in multi-agent systems and for evaluating agentic tool-calling. One study introduces AgentAudit…
-
New Qworld method generates question-specific LLM evaluation criteria
Researchers have introduced Qworld, a novel method for evaluating large language models (LLMs) by generating question-specific criteria. This approach addresses the limitations of static rubrics by creating detailed, co…
-
New HLE-Verified benchmark corrects errors in Humanity's Last Exam
Researchers have developed HLE-Verified, a revised version of the Humanity's Last Exam (HLE) benchmark designed to address concerns about noisy data and biased evaluations. The new benchmark employs a two-stage validati…
-
Sarvam 30B model performance metrics revealed across benchmarks
Sarvam AI has released its Sarvam 30B model, with performance metrics now available for several benchmarks. The model achieved 63.3% on GPQA, 7.5% on Humanity's Last Exam, and 19.2% on SciCode. Notably, it scored 0% on …
-
LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked
Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various…
-
OpenAI launches GPT-5.6 Sol with 14x faster "Ultrafast Mode" · 2 sources tracked
OpenAI has introduced an "Ultrafast Mode" for its GPT-5.6 Sol model, enabling speeds up to 14 times faster than previous iterations. This new service tier, available through the OpenAI API and powered by Cerebras, can d…
-
LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models
Independent benchmarks reveal varying performance across several large language models. Kimi K2 0905 achieved strong scores on GPQA and MMLU-Pro, while Qwen3 235B A22B also performed well on these metrics but struggled …
-
Meta's Muse Spark 1.2 shows rapid performance gains, rivals top AI models
Meta's latest foundational model, Muse Spark 1.2, has achieved high scores in third-party performance analyses, demonstrating rapid improvement since the Muse series' debut four months ago. The model notably surpassed G…
-
Nova Premier: Fast but Weak on Reasoning Benchmark
The Nova Premier model demonstrates impressive speed, processing 67.6 tokens per second. However, its reasoning capabilities are significantly weaker, achieving only 4.7% on the Humanity's Last Exam benchmark. This high…
-
Llama 3.1 Tulu3 405B shows performance gap on reasoning benchmarks
The Llama 3.1 Tulu3 405B model demonstrated a significant performance disparity across benchmarks, achieving 71.6% on MMLU-Pro while scoring only 3.5% on Humanity's Last Exam. This wide gap highlights the ongoing challe…
-
Thinking Machines releases Inkling-Small, outperforming larger predecessor
Thinking Machines Lab has launched Inkling-Small, a new open-weights multimodal model that prioritizes efficiency over sheer size. Despite being significantly smaller than its predecessor, Inkling, Inkling-Small demonst…
-
Study finds HLE benchmark measures general reasoning, not distinct LLM capabilities
A new study analyzing the Humanity's Last Exam (HLE) benchmark has found that its multiple-choice subset, comprising 428 items, primarily measures a single general reasoning factor rather than distinct subject-domain ca…
-
2026 LLM Benchmark: No Single Winner, Specialized Leaders Emerge · 1 source tracked
A comprehensive benchmark of 20 leading LLMs in 2026 reveals no single dominant model, but rather specialized leaders across different tasks. Claude Opus 5 leads the overall Artificial Analysis Intelligence Index, while…
-
Anthropic's Claude Opus 5 excels in bio/cyber tasks, bypassing Fable 5 restrictions
Anthropic has released Claude Opus 5, positioning it as a powerful tool for computational biology and cybersecurity tasks. While the "frontier" model, Fable 5, is heavily restricted in these domains, and Mythos 5 is gat…
-
Anthropic's new Opus model surpasses Fable 5 on key AI benchmarks
Anthropic has reportedly released a new Opus series model that outperforms its previous Fable 5 model on benchmarks like Humanity's Last Exam, agentic coding, and ARC-AGI. This development suggests significant advanceme…
-
LLM benchmark results reveal performance across multiple models · 9 sources tracked
A recent independent benchmark evaluation has revealed performance metrics for several large language models, including Kimi K2, Sarvam Maya, NVIDIA Nemotron 3 Super 120B, DeepSeek V3.2, Falcon H1R-7B, GLM-5.2, GLM-5.1,…