PulseAugur
EN
LIVE 21:56:49
ENTITY GDPval-AA

GDPval-AA

PulseAugur coverage of GDPval-AA — every cluster mentioning GDPval-AA across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
2
7 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
0
0 over 90d
TIER MIX · 90D
TOPICS
SENTIMENT · 30D

2 day(s) with sentiment data

RECENT · PAGE 1/1 · 7 TOTAL
  1. SIGNIFICANT · CL_166122 ·

    Anthropic's Claude Opus 5 ships with dynamic tool changes and improved benchmarks

    Anthropic has released Claude Opus 5, maintaining the price of Opus 4.8 while claiming near frontier intelligence and offering significant benchmark improvements. The release includes two beta features: the ability to d…

  2. FRONTIER RELEASE · CL_162004 ·

    Anthropic launches Claude Opus 5, challenging Fable 5 on benchmarks

    Anthropic has launched Claude Opus 5, a new frontier model that offers performance close to Claude Fable 5 at a reduced cost. Early evaluations show Opus 5 excelling in coding and knowledge work, setting new state-of-th…

  3. TOOL · CL_106672 ·

    GLM-5.2 leads open weights models on real-world agentic work benchmark · 2 sources tracked

    GLM-5.2 has emerged as the most popular new model on the Fireworks AI platform over the past week. This open-weights model has achieved the third overall position on the GDPval-AA benchmark, which evaluates performance …

  4. SIGNIFICANT · CL_95355 ·

    Fireworks AI offers Zhipu AI's GLM-5.2, top open-weights coding model

    Fireworks AI has announced that GLM-5.2 is now available on its inference platform, highlighting its performance as the top-ranked open-weights model for coding and third overall on the GDPval-AA benchmark. The model, d…

  5. SIGNIFICANT · CL_39378 ·

    Google DeepMind releases Gemini 3.5 Flash for faster agentic tasks

    Google DeepMind has launched Gemini 3.5 Flash, a new frontier intelligence model optimized for speed and agentic tasks. This model excels at complex, long-horizon tasks in coding and agent development, outperforming pre…

  6. FRONTIER RELEASE · CL_11237 ·

    X launches Grok 4.3 with improved agentic performance and lower price

    xAI has released Grok-4.3, a new iteration of its AI model, which offers improved agentic performance and a lower price point compared to its predecessor. The model achieved a significant increase of 321 ELO points on t…

  7. FRONTIER RELEASE · CL_07657 ·

    Xiaomi's MiMo-v2.5-Pro open-source model rivals top AI coding assistants

    Xiaomi has released MiMo-v2.5-Pro, an open-source coding-focused language model that demonstrates impressive capabilities in complex tasks. The model successfully completed a university-level compiler project in hours, …