GDPval-AA v2
PulseAugur coverage of GDPval-AA v2 — every cluster mentioning GDPval-AA v2 across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Anthropic's Claude Opus 5 leads agentic index, Qwen3.8 Max close behind · 1 source tracked
Artificial Analysis's Agentic Index shows Anthropic's Claude Opus 5 leading, with Alibaba Group's Qwen3.8 Max closely following. While some reports incorrectly declared Qwen the top model, the index actually places Clau…
-
Meta's Muse Spark 1.2 shows rapid performance gains, rivals top AI models
Meta's latest foundational model, Muse Spark 1.2, has achieved high scores in third-party performance analyses, demonstrating rapid improvement since the Muse series' debut four months ago. The model notably surpassed G…
-
Google launches cheaper Gemini Flash models, prioritizing cost over peak performance
Google has released three new Gemini Flash models—3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber—focused on cost-efficiency for large-scale AI agents. The 3.6 Flash model offers reduced output token consumption and lowe…
-
Anthropic releases Claude Sonnet 5 with enhanced agentic capabilities
Anthropic has released Claude Sonnet 5, an updated mid-tier model that significantly improves agentic capabilities and performance over its predecessor, Sonnet 4.6. This new model demonstrates enhanced abilities in plan…
-
Claude Fable 5 leads AI performance leaderboard but is most expensive
A new evaluation called gdpval-aa v2 measures AI model performance on real-world tasks, using an Elo rating system anchored to a human baseline. Anthropic's Claude Fable 5, Sonnet 5, and Opus 4.8 models secured the top …