GDPval-AA v2
PulseAugur coverage of GDPval-AA v2 — every cluster mentioning GDPval-AA v2 across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Moonshot launches Kimi K3 with 2.8T parameters and 1M context window
Moonshot has launched its Kimi K3 model, a 2.8-trillion-parameter Mixture-of-Experts model with a context window of over 1 million tokens. The model features a new Kimi Delta Attention mechanism, which combines linear a…
-
Artificial Analysis Intelligence Index v4.2 released, Anthropic leads rankings
Artificial Analysis has released version 4.2 of its Intelligence Index, introducing new evaluations like AA-Briefcase for agentic knowledge work and GDP.pdf for long-context document reasoning. This update increases the…
-
Anthropic's Claude Opus 5 leads agentic index, Qwen3.8 Max close behind · 1 source tracked
Artificial Analysis's Agentic Index shows Anthropic's Claude Opus 5 leading, with Alibaba Group's Qwen3.8 Max closely following. While some reports incorrectly declared Qwen the top model, the index actually places Clau…
-
Meta's Muse Spark 1.2 shows rapid performance gains, rivals top AI models
Meta's latest foundational model, Muse Spark 1.2, has achieved high scores in third-party performance analyses, demonstrating rapid improvement since the Muse series' debut four months ago. The model notably surpassed G…
-
Google launches cheaper Gemini Flash models, prioritizing cost over peak performance
Google has released three new Gemini Flash models—3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—focused on cost-efficiency for large-scale AI agents. The 3.6 Flash model offers reduced output token consumption and lowe…
-
Anthropic releases Claude Sonnet 5 with enhanced agentic capabilities
Anthropic has released Claude Sonnet 5, an updated mid-tier model that significantly improves agentic capabilities and performance over its predecessor, Sonnet 4.6. This new model demonstrates enhanced abilities in plan…
-
Claude Fable 5 leads AI performance leaderboard but is most expensive
A new evaluation called gdpval-aa v2 measures AI model performance on real-world tasks, using an Elo rating system anchored to a human baseline. Anthropic's Claude Fable 5, Sonnet 5, and Opus 4.8 models secured the top …