LiveBench
PulseAugur coverage of LiveBench — every cluster mentioning LiveBench across labs, papers, and developer communities, ranked by signal.
-
New paper proposes Bayesian audits for AI evaluation archives
A new paper proposes a Bayesian inference framework to audit public archives of frontier AI evaluations. The research highlights how selective reporting and benchmark revisions can distort the perception of AI progress,…
-
Anthropic's Fable 5 lags Gemini 3.1 on LiveBench benchmark
A new benchmark evaluation on LiveBench shows Fable 5 performing below Gemini 3.1. The results raise questions about the benchmark's accuracy or Anthropic's evaluation methodology. This performance dip for Fable 5, a mo…
-
AI benchmarks criticized as useless due to over-optimization and contamination
The author argues that current AI model benchmarks are becoming increasingly useless due to several factors. They contend that models are being over-optimized for these specific tests, leading to a disconnect between be…