Humanity's Last Exam
PulseAugur coverage of Humanity's Last Exam — every cluster mentioning Humanity's Last Exam across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
Anthropic's new Opus model surpasses Fable 5 on key AI benchmarks
Anthropic has reportedly released a new Opus series model that outperforms its previous Fable 5 model on benchmarks like Humanity's Last Exam, agentic coding, and ARC-AGI. This development suggests significant advanceme…
-
GLM-4.7-Flash model shows performance metrics across benchmarks
The GLM-4.7-Flash model has demonstrated specific performance metrics across several benchmarks, including GPQA, Humanity's Last Exam, Long Context Reasoning, and SciCode. The model achieved 45.2% on GPQA and 25.5% on S…
-
Developer builds quiz to test human vs. AI on expert-level exam
A developer has created a web-based quiz called "Humans vs. Humanity's Last Exam" that pits human players against advanced AI models on challenging academic questions. The quiz utilizes the "Humanity's Last Exam" datase…
-
DBRX Instruct and Mistral Medium 3 benchmark results revealed
Independent benchmarks reveal performance metrics for two large language models. DBRX Instruct achieved scores of 33.1% on GPQA, 39.7% on MMLU-Pro, 6.6% on Humanity's Last Exam, and 9.3% on LiveCodeBench. Mistral Medium…
-
New AI agents tackle deep research and misleading web data · 4 sources tracked
Researchers have introduced AREX, a new family of recursively self-improving agents designed for deep research tasks. AREX alternates between research and self-improvement loops, using an autonomous context-update tool …
-
Mira Murati's Thinking Machines releases open-source Inkling model
Thinking Machines, co-founded by former OpenAI executive Mira Murati, has released its first model, Inkling. Unlike many frontier models, Inkling does not aim to top leaderboards, scoring lower than models like Claude F…
-
Open-source LLMs show strong benchmark performance across multiple metrics · 4 sources tracked
Several open-source AI models have demonstrated strong performance on various benchmarks, according to independent measurements. Mi:dm K 2.5 Pro achieved 70.1% on GPQA and 80.9% on MMLU-Pro, while MiMo-V2-Flash showed 8…
-
AI Benchmark 'Humanity's Last Exam' Criticized as Distraction
The article "Humanity's Last Exam" critiques the AI evaluation benchmark, exploring its origins and the varied expert opinions surrounding its creation. It suggests that the benchmark may serve as a distraction from mor…
-
OpenClaw AI agent framework matures, gains wider adoption
OpenClaw, an open-source AI agent framework, has matured significantly since its release a few months ago, evolving from a niche tool to a widely adopted local-first assistant. It can now execute real-world tasks by con…
-
Sakana Fugu orchestrator models combine LLMs for collective intelligence
Researchers have developed Sakana Fugu, a family of orchestrator models designed to combine the specialized capabilities of multiple Large Language Models (LLMs) into a collectively intelligent system. These models act …
-
Perplexity Integrates Deep Research with Multi-Model Orchestration System
Perplexity has integrated its Deep Research feature into its Computer orchestration system, enhancing its ability to break down complex questions into subtasks. These subtasks are then routed across more than 20 differe…
-
Andon Labs stress-tests AI agents in real-world business scenarios
Andon Labs is developing novel real-world evaluations for AI systems, moving beyond traditional benchmarks to assess model behavior in complex scenarios. Their "Vending-Bench" and "Luna" projects, which involve AI-run p…
-
Google's Gemini 3.5 Flash outperforms 3.1 Pro on coding and agents
Google's Gemini 3.5 Flash model has surpassed its predecessor, Gemini 3.1 Pro, on several key benchmarks, particularly in coding and agentic tasks. This new tier offers a significant cost reduction of 40% and approximat…
-
LLMs learn to actively seek external info for better task adaptation
Researchers have developed a new method for adapting large language models (LLMs) by enabling them to actively seek information from external sources like Wikipedia and web browsers. This approach, termed "active inform…
-
New RSE strategy recycles LLM search experience for efficient test-time scaling
Researchers have introduced Recycling Search Experience (RSE), a novel method to improve the efficiency of test-time scaling for large language models. RSE transforms test-time search from isolated trials into a cumulat…
-
OpenSearch-VL offers open recipe for advanced multimodal search agents
Researchers have developed OpenSearch-VL, a novel, fully open-source recipe for training advanced multimodal deep search agents. This approach utilizes a curated pipeline for high-quality training data, a diverse tool e…
-
Xiaomi's MiMo-v2.5-Pro open-source model rivals top AI coding assistants
Xiaomi has released MiMo-v2.5-Pro, an open-source coding-focused language model that demonstrates impressive capabilities in complex tasks. The model successfully completed a university-level compiler project in hours, …
-
MTRouter cuts LLM costs by 58% on ScienceWorld, 43% on HLE
Researchers have developed MTRouter, a novel system designed to optimize the cost of multi-turn interactions with large language models. By jointly embedding interaction history and candidate models, MTRouter learns to …
-
Google Gemini API adds Deep Research updates with MCP and chart generation
Google has released two significant updates to its Gemini API, enhancing its Deep Research capabilities. These updates introduce improved quality, support for MCP, and native generation of charts and infographics. The G…
-
Google upgrades Gemini 3 Deep Think for science and engineering
Google has released an upgraded version of Gemini 3 Deep Think, a specialized reasoning mode designed for complex scientific, research, and engineering challenges. This new iteration is available to Google AI Ultra subs…