MMLU-Pro
PulseAugur coverage of MMLU-Pro — every cluster mentioning MMLU-Pro across labs, papers, and developer communities, ranked by signal.
7 day(s) with sentiment data
-
Qwen3 Omni 30B A3B Instruct shows mixed benchmark results, excelling in some areas but failing in long-context reasoning
The Qwen3 Omni 30B A3B Instruct model has demonstrated strong performance on specific benchmarks, achieving 62% on GPQA and 72.5% on MMLU-Pro. However, it showed no capability in long-context reasoning. The model operat…
-
Research finds truncation samplers offer limited benefit at typical LLM temperatures
A new research paper explores the impact of temperature sampling and truncation methods on large language model performance. The study found that while truncation samplers like top-p and min-p are often associated with …
-
New framework audits test-time scaling for video world models
A new research paper introduces the Compute-Value Audit (CVA) framework to evaluate the effectiveness of test-time scaling (TTS) in video world models. The study found that while increasing sampling can improve candidat…
-
Seed-OSS-36B-Instruct achieves high cost-efficiency in benchmarks
Seed-OSS-36B-Instruct has achieved a cost-efficiency of 40.3 intelligence points per dollar, according to independent measurements. The model scored 72.6% on the GPQA benchmark and 81.5% on MMLU-Pro. These results sugge…
-
Mistral Small 3.2 released with enhanced function calling and 128K context
Mistral AI has released Mistral Small 3.2, an open-weight model featuring improved function calling and a 128K context window. This update enhances the model's ability to handle tool distinctions and provides cleaner JS…
-
MoE models achieve shorter reasoning with inference-time routing tweaks
A new research paper introduces a method to reduce the number of reasoning tokens and latency in Mixture-of-Experts (MoE) models without requiring retraining. By adjusting the router at inference time to allocate more e…
-
BenchMIRT method reveals what LLM benchmarks truly measure · 2 sources tracked
Researchers have introduced BenchMIRT, a novel methodology designed to dissect the performance of large language models (LLMs) on benchmarks by analyzing individual prompts. This approach, inspired by Item Response Theo…
-
LLM Benchmark Results: Molmo, Qwen3, Hermes, and Mistral Performance Revealed
Independent benchmarks reveal varying performance across several large language models. Molmo 7B-D shows low scores on GPQA and MMLU-Pro, while Qwen3 Omni 30B A3B and Hermes 4 - Llama-3.1 405B demonstrate significantly …
-
LFM2 8B A1B model shows mixed benchmark results
The LFM2 8B A1B model achieved a score of 50.5% on the MMLU-Pro benchmark. However, its performance on Long Context Reasoning was notably low, scoring only 1%. This disparity suggests potential issues with the evaluatio…
-
LLM judges in multi-agent systems face reliability issues, new research suggests
Multiple research papers explore the limitations and potential improvements of using Large Language Models (LLMs) as judges in multi-agent systems and for evaluating agentic tool-calling. One study introduces AgentAudit…
-
OmnisRouter benchmarked near bottom due to high cost, not routing skill
The developer of OmnisRouter, a tool designed to route LLM requests to the most cost-effective model, has shared benchmark results that place their router near the bottom. Despite initial bugs in the testing setup, the …
-
Alibaba previews Qwen4 architecture with cost-efficient Qwen3.8-Flash-Next model
Alibaba's Qwen team has released Qwen3.8-Flash-Next, an open-weight multimodal MoE model that previews the architecture for the upcoming Qwen4. This new model boasts significant cost-efficiency, activating only 6B param…
-
Developer tests reveal Qwen3 variants perform differently than benchmarks suggest
A developer compared the performance of Qwen2.5 and Qwen3 models using a custom script with 40 specific prompts related to ticket classification. While Qwen3's published benchmarks indicated broad improvements, the deve…
-
New Tiel-Coder 35B model excels at coding and long conversations
A new open-source model, Tiel-Coder-35B-A3B, has been released, optimized for coding tasks and long conversations. It achieves strong performance on the SWE-bench-Live benchmark, fixing 12 out of 25 problems, which is c…
-
LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked
Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various…
-
Mi:dm K 2.5 Pro shows strong MMLU-Pro performance but struggles with long context
The Mi:dm K 2.5 Pro model has achieved an 80.9% score on the MMLU-Pro benchmark. However, its performance on long-context reasoning tasks is significantly lower, registering only 10%. These results were independently me…
-
New framework detects hidden behavioral entanglement in LLMs
Researchers have developed a new statistical framework to detect and quantify behavioral entanglement among large language models (LLMs). This framework uses information-theoretic metrics, specifically a Difficulty-Weig…
-
Lightweight fine-tuning prunes MoE models, reducing size and latency
Researchers have developed a method to prune experts in Mixture-of-Experts (MoE) models using lightweight fine-tuning techniques. By applying parameter-efficient adapters like LoRA, they can identify and remove less cri…
-
LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models
Independent benchmarks reveal varying performance across several large language models. Kimi K2 0905 achieved strong scores on GPQA and MMLU-Pro, while Qwen3 235B A22B also performed well on these metrics but struggled …
-
HELENA framework enhances multi-agent systems with novel coordination
Researchers have introduced HELENA, a novel multi-agent system framework designed to enhance analytical capacity by integrating diverse reasoning paths while mitigating noise. HELENA constructs a composite graph from co…