SummEval
PulseAugur coverage of SummEval — every cluster mentioning SummEval across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
LLM-as-a-Judge: Verbalized Confidence Outperforms Log-Probabilities on New Models
A new arXiv paper proposes a shift in how Large Language Models (LLMs) are used as judges, suggesting that verbalized confidence is now a more robust scoring mechanism than log-probabilities for post-2025 proprietary mo…
-
LLM judge panel calibration framework introduced
Researchers have developed a framework called Finite-Calibration Panel Selection to determine the optimal calibration strategy for LLM judge panels. This method helps decide whether to use low-dimensional stackers or jo…
-
New LLM evaluation methods tackle alignment and bias
Researchers are developing new methods to evaluate and improve the alignment and interpretability of large language models (LLMs). Google Research has introduced a framework that adapts psychological assessments to quan…