PulseAugur
EN
LIVE 13:17:11
ENTITY GPQA: A Graduate-Level Google-Proof Q&A Benchmark

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

PulseAugur coverage of GPQA: A Graduate-Level Google-Proof Q&A Benchmark — every cluster mentioning GPQA: A Graduate-Level Google-Proof Q&A Benchmark across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
15
38 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
8
21 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

10 day(s) with sentiment data

RECENT · PAGE 1/3 · 58 TOTAL
  1. RESEARCH · CL_258923 ·

    DeepSeek V4 Pro and Nova 2.0 Lite benchmarks revealed · 2 sources tracked

    Independent benchmarks reveal the performance of two large language models, DeepSeek V4 Pro and Nova 2.0 Lite. DeepSeek V4 Pro achieved high scores in reasoning-focused tasks, with 92.8% on GPQA and 80.3% on Long Contex…

  2. TOOL · CL_256749 ·

    Qwen3 Omni 30B A3B Instruct shows mixed benchmark results, excelling in some areas but failing in long-context reasoning

    The Qwen3 Omni 30B A3B Instruct model has demonstrated strong performance on specific benchmarks, achieving 62% on GPQA and 72.5% on MMLU-Pro. However, it showed no capability in long-context reasoning. The model operat…

  3. TOOL · CL_250827 ·

    DeepSeek-V4 Flash 0731 achieves strong GPQA and HLE scores

    DeepSeek-V4 Flash 0731 has achieved a score of 90.8% on the GPQA benchmark and 38.6% on HLE. The model also demonstrated a speed of 233.8 tokens per second. Notably, it offers a high intelligence-to-cost ratio, providin…

  4. TOOL · CL_243485 ·

    Seed-OSS-36B-Instruct achieves high cost-efficiency in benchmarks

    Seed-OSS-36B-Instruct has achieved a cost-efficiency of 40.3 intelligence points per dollar, according to independent measurements. The model scored 72.6% on the GPQA benchmark and 81.5% on MMLU-Pro. These results sugge…

  5. TOOL · CL_240266 ·

    Solar Pro 4 achieves 89.1% on GPQA, shows strong efficiency

    Solar Pro 4 has achieved an 89.1% score on the GPQA benchmark, demonstrating strong performance in graduate-level question answering. The model also exhibited an efficiency of 59.1 tokens per second with 62.3 intelligen…

  6. TOOL · CL_237569 ·

    GLM-5.1 (Non-reasoning) achieves 83.9% on GPQA benchmark

    A new benchmark evaluation shows that GLM-5.1 (Non-reasoning) achieved an 83.9% score on the GPQA benchmark. This performance was measured independently and highlights the model's efficiency, delivering 13.3 intelligenc…

  7. TOOL · CL_232771 ·

    Qwen3.5 2B shows mixed results in independent benchmarks

    Independent benchmarks reveal that Qwen3.5 2B, a non-reasoning model, achieves 43.8% on the GPQA benchmark. However, its performance significantly drops on more complex tasks, scoring only 5% on HLE, 15% on Long Context…

  8. TOOL · CL_232335 ·

    Qwen3.6 and Qwen3.5 show similar inference speeds, with gains in agentic tasks

    A recent benchmark comparison of Qwen3.6 and Qwen3.5 models revealed that their inference speeds on a GeForce RTX 4070 were nearly identical, contrary to initial findings that suggested a significant slowdown. This disc…

  9. TOOL · CL_231543 ·

    New research suggests RL enhances language model sampling efficiency, not new reasoning

    A new research paper explores how reinforcement learning (RL) impacts language model reasoning, specifically whether it introduces new reasoning capabilities or enhances the sampling of existing ones. The study introduc…

  10. RESEARCH · CL_230863 ·

    BenchMIRT method reveals what LLM benchmarks truly measure · 2 sources tracked

    Researchers have introduced BenchMIRT, a novel methodology designed to dissect the performance of large language models (LLMs) on benchmarks by analyzing individual prompts. This approach, inspired by Item Response Theo…

  11. RESEARCH · CL_227560 ·

    New PRACT-120 benchmark aims to evaluate AI chatbots holistically

    A new benchmark called PRACT-120 has been proposed to evaluate AI chatbots more comprehensively than existing tests like MMLU or GPQA. The benchmark aims to assess not just the core model's capabilities but also the int…

  12. RESEARCH · CL_224467 ·

    LLM Benchmark Results: Molmo, Qwen3, Hermes, and Mistral Performance Revealed

    Independent benchmarks reveal varying performance across several large language models. Molmo 7B-D shows low scores on GPQA and MMLU-Pro, while Qwen3 Omni 30B A3B and Hermes 4 - Llama-3.1 405B demonstrate significantly …

  13. TOOL · CL_223254 ·

    New TRACES framework enables cost-efficient early stopping for LLM reasoning

    Researchers have introduced TRACES, a new framework designed to tag reasoning steps in Language Reasoning Models (LRMs) to enable adaptive and cost-efficient early stopping. This method monitors reasoning behaviors duri…

  14. TOOL · CL_222734 ·

    Kimi K2.5 achieves strong benchmark scores with competitive pricing

    Kimi K2.5 has achieved notable performance on several benchmarks, including GPQA, HLE, Long Context, and SciCode. The model offers competitive pricing at 30 integer points per dollar across these evaluations. These resu…

  15. RESEARCH · CL_228971 ·

    LLM judges in multi-agent systems face reliability issues, new research suggests

    Multiple research papers explore the limitations and potential improvements of using Large Language Models (LLMs) as judges in multi-agent systems and for evaluating agentic tool-calling. One study introduces AgentAudit…

  16. TOOL · CL_218073 ·

    New GIM benchmark evaluates LLMs on integrated cognitive tasks

    Researchers have introduced the Grounded Integration Measure (GIM), a new benchmark designed to evaluate large language models (LLMs) by assessing their ability to integrate multiple cognitive operations. Unlike benchma…

  17. TOOL · CL_210684 ·

    Qwen3.5 35B A3B model achieves high GPQA score and cost-efficiency

    The Qwen3.5 35B A3B model has demonstrated impressive performance, achieving an 81.9% score on the GPQA benchmark. It also offers a high intelligence-to-cost ratio, delivering 35.3 intelligence points per dollar. The mo…

  18. SIGNIFICANT · CL_204942 ·

    Anthropic's Claude 3.5 Sonnet overtakes Opus in benchmarks, becomes default for builders

    Anthropic has released Claude 3.5 Sonnet, a new mid-tier model that surpasses its previous flagship, Claude 3 Opus, in reasoning and coding benchmarks. This release offers a significant improvement in the cost-performan…

  19. TOOL · CL_203015 ·

    Sarvam 30B model performance metrics revealed across benchmarks

    Sarvam AI has released its Sarvam 30B model, with performance metrics now available for several benchmarks. The model achieved 63.3% on GPQA, 7.5% on Humanity's Last Exam, and 19.2% on SciCode. Notably, it scored 0% on …

  20. RESEARCH · CL_201933 ·

    LLM performance benchmarks released for Llama, GLM, and Mistral models · 4 sources tracked

    Independent benchmarks reveal performance metrics for several large language models, including Llama 3.2 Instruct 90B, GLM-4.7-Flash, Mistral Large 2, and Llama 3.1 Instruct 8B. The data highlights scores across various…