Zheng et al. reply
PulseAugur coverage of Zheng et al. reply — every cluster mentioning Zheng et al. reply across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
AI evaluation scores are flawed, focusing on models over graders
A recent analysis highlights a critical flaw in AI model evaluation: the focus is overwhelmingly on the model's performance, while the reliability of the evaluation instrument itself is often neglected. An anecdote illu…
-
LLM Agreement Weak Proxy for Accuracy, Study Finds
A new arXiv paper investigates the reliability of using agreement among Large Language Models (LLMs) as a proxy for correctness. The study, which involved 53 different LLM runners and 265,000 samples, found that while a…
-
LLM judges show inconsistency and bias, requiring new evaluation methods
Large language models used as judges in automated evaluation systems can exhibit inconsistencies, leading to unreliable results. Factors such as sampling temperature, model version drift, prompt ambiguity, and tie-break…