PulseAugur
EN
LIVE 15:03:44
ENTITY MMLU-Pro

MMLU-Pro

PulseAugur coverage of MMLU-Pro — every cluster mentioning MMLU-Pro across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
19
36 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
11
24 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

11 day(s) with sentiment data

RECENT · PAGE 1/2 · 36 TOTAL
  1. TOOL · CL_193611 ·

    New framework detects hidden behavioral entanglement in LLMs

    Researchers have developed a new statistical framework to detect and quantify behavioral entanglement among large language models (LLMs). This framework uses information-theoretic metrics, specifically a Difficulty-Weig…

  2. TOOL · CL_193458 ·

    Lightweight fine-tuning prunes MoE models, reducing size and latency

    Researchers have developed a method to prune experts in Mixture-of-Experts (MoE) models using lightweight fine-tuning techniques. By applying parameter-efficient adapters like LoRA, they can identify and remove less cri…

  3. RESEARCH · CL_188546 ·

    LLM benchmarks show mixed results for Kimi, Qwen3, and Exaone models

    Independent benchmarks reveal varying performance across several large language models. Kimi K2 0905 achieved strong scores on GPQA and MMLU-Pro, while Qwen3 235B A22B also performed well on these metrics but struggled …

  4. TOOL · CL_184946 ·

    HELENA framework enhances multi-agent systems with novel coordination

    Researchers have introduced HELENA, a novel multi-agent system framework designed to enhance analytical capacity by integrating diverse reasoning paths while mitigating noise. HELENA constructs a composite graph from co…

  5. TOOL · CL_182335 ·

    LLM prompt engineering: Specificity beats politeness, research shows

    Prompt engineering best practices are evolving, with new research suggesting that elements like politeness and persona do not reliably improve LLM performance. Studies from The Wharton School and findings from EMNLP 202…

  6. TOOL · CL_182150 ·

    Quantiles enables no-code AI evaluations with Hugging Face datasets

    Quantiles has introduced a new feature allowing users to create AI evaluations without writing code, utilizing a configuration-based approach. This system supports deterministic scoring styles like exact match and multi…

  7. TOOL · CL_177381 ·

    DeepSeek V4 flash version shows strong performance on MMLU-Pro, GPQA Diamond

    DeepSeek V4 has released a new "flash" version, reportedly achieving impressive scores on benchmarks like MMLU-Pro, GPQA Diamond, and TruthfulQA. The model is noted for its strong performance relative to its size, with …

  8. TOOL · CL_176863 ·

    Llama 3.1 Tulu3 405B shows performance gap on reasoning benchmarks

    The Llama 3.1 Tulu3 405B model demonstrated a significant performance disparity across benchmarks, achieving 71.6% on MMLU-Pro while scoring only 3.5% on Humanity's Last Exam. This wide gap highlights the ongoing challe…

  9. TOOL · CL_167276 ·

    New HG-CRC framework enhances LLM risk control across subgroups

    Researchers have developed a new framework called Hierarchical Group-Conditional Conformal Risk Control (HG-CRC) to improve the reliability of large language models. This method ensures that risk guarantees are met not …

  10. TOOL · CL_167183 ·

    New framework optimizes multi-LLM question planning for cost and quality

    Researchers have developed OPTI-Q, a new framework designed to optimize question planning for multiple large language models (LLMs). This system uses a database-inspired, cost-based optimizer to create execution plans t…

  11. RESEARCH · CL_161409 ·

    LLM benchmark results reveal performance across multiple models · 9 sources tracked

    A recent independent benchmark evaluation has revealed performance metrics for several large language models, including Kimi K2, Sarvam Maya, NVIDIA Nemotron 3 Super 120B, DeepSeek V3.2, Falcon H1R-7B, GLM-5.2, GLM-5.1,…

  12. TOOL · CL_160746 ·

    LLM diversity metrics may not measure diversity, study finds

    A new research paper published on arXiv questions the effectiveness of common diversity metrics used in Large Language Model (LLM) ensembles. The study found that these metrics often correlate more with the models' over…

  13. TOOL · CL_160721 ·

    New method estimates LLM trust using expert judgment and calibration

    Researchers have developed a novel method for estimating trust in multi-Large Language Model (LLM) systems by adapting structured expert judgment techniques. This approach, termed uncertainty-aware trust estimation, use…

  14. TOOL · CL_158645 ·

    Solar Open 2 language model boasts 1M-token context window and strong agentic skills

    Researchers have introduced Solar Open 2, a 250 billion parameter Mixture-of-Experts language model designed for long-horizon agentic tasks. This model features a 1 million token context window achieved through a hybrid…

  15. RESEARCH · CL_152721 ·

    DBRX Instruct and Mistral Medium 3 benchmark results revealed

    Independent benchmarks reveal performance metrics for two large language models. DBRX Instruct achieved scores of 33.1% on GPQA, 39.7% on MMLU-Pro, 6.6% on Humanity's Last Exam, and 9.3% on LiveCodeBench. Mistral Medium…

  16. RESEARCH · CL_150653 ·

    Chinese Open-Source LLMs Offer Major Cost Savings Over GPT-5

    Open-source Large Language Models from China are emerging as a cost-effective alternative to models like GPT-5, with some startups reporting significant savings. These models leverage efficient Mixture-of-Experts (MoE) …

  17. TOOL · CL_147965 ·

    New Step-Tagging framework enhances control over Language Reasoning Models

    Researchers have introduced a new framework called Step-Tagging to better control the generation process of Language Reasoning Models (LRMs). This framework uses a lightweight sentence classifier to annotate reasoning s…

  18. RESEARCH · CL_137517 ·

    Open-source LLMs show strong benchmark performance across multiple metrics · 4 sources tracked

    Several open-source AI models have demonstrated strong performance on various benchmarks, according to independent measurements. Mi:dm K 2.5 Pro achieved 70.1% on GPQA and 80.9% on MMLU-Pro, while MiMo-V2-Flash showed 8…

  19. TOOL · CL_120530 ·

    NVIDIA quantizes Mistral Medium 3.5 for reduced GPU memory usage

    NVIDIA has quantized the Mistral Medium 3.5 (128B) model using its Model Optimizer v0.44.0 and the NVFP4 quantization method. This process significantly reduces GPU memory requirements with negligible loss in accuracy, …

  20. RESEARCH · CL_119406 ·

    New 'LearnStop' method optimizes reasoning model stopping points

    Researchers have developed a new method called LearnStop to optimize when reasoning language models should stop processing an instance. This technique analyzes multiple features like answer confidence, entropy, and stab…