PulseAugur
EN
LIVE 10:49:24
ENTITY GPQA Diamond

GPQA Diamond

PulseAugur coverage of GPQA Diamond — every cluster mentioning GPQA Diamond across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
15
44 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
6
24 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

10 day(s) with sentiment data

RECENT · PAGE 1/3 · 44 TOTAL
  1. TOOL · CL_182335 ·

    LLM prompt engineering: Specificity beats politeness, research shows

    Prompt engineering best practices are evolving, with new research suggesting that elements like politeness and persona do not reliably improve LLM performance. Studies from The Wharton School and findings from EMNLP 202…

  2. SIGNIFICANT · CL_182151 ·

    Google DeepMind's DiffusionGemma achieves 1500 tokens/sec via discrete diffusion

    Google DeepMind has released DiffusionGemma, an open-weight language model that utilizes discrete diffusion for text generation, offering significantly faster output speeds compared to traditional autoregressive models.…

  3. TOOL · CL_178393 ·

    New Metanym Game benchmark evaluates LLM structural intelligence

    Researchers have introduced the Metanym Game, a novel benchmark designed to evaluate the structural intelligence of Large Language Models (LLMs). This game operates as a self-contained, self-consistent system where LLMs…

  4. SIGNIFICANT · CL_178669 ·

    Alibaba launches Qwen3.8, enhancing coding and office AI capabilities · 2 sources tracked

    Alibaba has officially launched its new flagship large language model, Qwen3.8, boasting a total parameter count of 2.4 trillion. This advanced model demonstrates significant improvements in programming and professional…

  5. TOOL · CL_177381 ·

    DeepSeek V4 flash version shows strong performance on MMLU-Pro, GPQA Diamond

    DeepSeek V4 has released a new "flash" version, reportedly achieving impressive scores on benchmarks like MMLU-Pro, GPQA Diamond, and TruthfulQA. The model is noted for its strong performance relative to its size, with …

  6. FRONTIER RELEASE · CL_175302 ·

    Thinking Machines releases Inkling-Small, outperforming larger predecessor

    Thinking Machines Lab has launched Inkling-Small, a new open-weights multimodal model that prioritizes efficiency over sheer size. Despite being significantly smaller than its predecessor, Inkling, Inkling-Small demonst…

  7. TOOL · CL_174388 ·

    Darwin AI model family achieves 90.9% on GPQA Diamond via evolutionary merging

    The Darwin AI model family achieves a 90.9% score on the GPQA Diamond benchmark by using evolutionary merging of existing open-weight models, rather than traditional pretraining. This approach, which combines models lik…

  8. RESEARCH · CL_173812 ·

    Korean startup's Darwin-398B-JGOS model ranks 3rd globally on GPQA Diamond

    A South Korean startup, VIDRAFT, has developed a language model named Darwin-398B-JGOS that achieved the top rank among Korean models on the GPQA Diamond benchmark. This model, reportedly trained on approximately 24 GPU…

  9. COMMENTARY · CL_172646 ·

    2026 LLM Benchmark: No Single Winner, Specialized Leaders Emerge · 1 source tracked

    A comprehensive benchmark of 20 leading LLMs in 2026 reveals no single dominant model, but rather specialized leaders across different tasks. Claude Opus 5 leads the overall Artificial Analysis Intelligence Index, while…

  10. RESEARCH · CL_170794 ·

    Korean startup VIDRAFT achieves AI breakthroughs with 24 GPUs

    South Korean startup VIDRAFT has achieved significant results using a modest setup of approximately 24 GPUs, challenging the dominance of Big Tech's large-scale operations. The company's Darwin models have topped the Ko…

  11. TOOL · CL_154229 ·

    New diagnostic tool assesses LLM test-time collaboration effectiveness

    Researchers have developed a new diagnostic framework to evaluate the effectiveness of test-time collaboration techniques for large language models. This framework, called the "Fixed-Pool Diagnostic," decomposes the gai…

  12. RESEARCH · CL_153903 ·

    New law quantifies LLM ensemble diversity uplift · 2 sources tracked

    Researchers have developed a formal law to quantify the performance uplift gained from using diverse large language model (LLM) ensembles. This law decomposes ensemble lift into "rescue" and "damage" components, providi…

  13. SIGNIFICANT · CL_149118 ·

    Mira Murati's Thinking Machines releases open-source Inkling model

    Thinking Machines, co-founded by former OpenAI executive Mira Murati, has released its first model, Inkling. Unlike many frontier models, Inkling does not aim to top leaderboards, scoring lower than models like Claude F…

  14. SIGNIFICANT · CL_147230 ·

    Moonshot AI releases Kimi K3 with 1M context, open weights, and competitive pricing

    Moonshot AI has released Kimi K3, an open-weights model with approximately 2.8 trillion parameters and a 1 million token context window. The model achieved strong performance on benchmarks like Terminal-Bench 2.0, Front…

  15. TOOL · CL_139122 ·

    VIDRAFT ships dual LLM serving engines for GPU throughput and CPU reach

    VIDRAFT has developed two distinct serving engines for large language models, addressing separate optimization targets. VKAE is a kernel-level acceleration engine designed to maximize throughput on GPUs, achieving up to…

  16. SIGNIFICANT · CL_138963 ·

    Unisound U2: 266B model achieves high benchmark score at low cost

    Unisound, a Chinese company primarily known for speech AI, has released a 266 billion parameter large language model named Unisound U2. This model achieved a score of 87.9% on the GPQA Diamond benchmark, a test designed…

  17. RESEARCH · CL_138139 ·

    Tencent and VIDRAFT showcase sparse MoE models with reduced active parameters

    Tencent has released Hy3, a 295-billion-parameter Mixture-of-Experts (MoE) model that utilizes only 21 billion active parameters per forward pass, significantly reducing inference costs. This MoE architecture, featuring…

  18. TOOL · CL_135313 ·

    LLM Agreement Weak Proxy for Accuracy, Study Finds

    A new arXiv paper investigates the reliability of using agreement among Large Language Models (LLMs) as a proxy for correctness. The study, which involved 53 different LLM runners and 265,000 samples, found that while a…

  19. SIGNIFICANT · CL_129683 ·

    Tencent releases Hy3, an open 295B MoE model with 256K context

    Tencent has released Hy3, an open-source 295 billion parameter Mixture-of-Experts (MoE) model designed for complex reasoning, agentic workflows, and long-context tasks. The model activates only 21 billion parameters per…

  20. COMMENTARY · CL_129702 ·

    AI benchmark charts: How to spot saturation and contamination

    A guide to interpreting AI benchmark charts, particularly for 2026 models, highlights the limitations and potential for misrepresentation in common evaluations. Benchmarks like SWE-bench Pro are introduced to combat dat…