LiveBench
PulseAugur coverage of LiveBench — every cluster mentioning LiveBench across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
DeepSeek V4.1 drastically undercuts Fable 5.1 on coding task costs
DeepSeek V4.1 has demonstrated a significant cost advantage over Fable 5.1 when performing coding tasks. According to data from LiveBench, DeepSeek V4.1 completed a coding task for $0.04, whereas Fable 5.1 incurred a co…
-
SCX Router unveiled for zero-shot LLM selection
Researchers have developed the SCX Router, a novel system designed to intelligently select the most suitable large language model (LLM) for a given task. This lightweight router, based on GLiClass and utilizing a Qwen3 …
-
New paper proposes Bayesian audits for AI evaluation archives
A new paper proposes a Bayesian inference framework to audit public archives of frontier AI evaluations. The research highlights how selective reporting and benchmark revisions can distort the perception of AI progress,…
-
Anthropic's Fable 5 lags Gemini 3.1 on LiveBench benchmark
A new benchmark evaluation on LiveBench shows Fable 5 performing below Gemini 3.1. The results raise questions about the benchmark's accuracy or Anthropic's evaluation methodology. This performance dip for Fable 5, a mo…
-
AI benchmarks criticized as useless due to over-optimization and contamination
The author argues that current AI model benchmarks are becoming increasingly useless due to several factors. They contend that models are being over-optimized for these specific tests, leading to a disconnect between be…