Bradley-Terry-Luce model
PulseAugur coverage of Bradley-Terry-Luce model — every cluster mentioning Bradley-Terry-Luce model across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New research frames LLM evaluation as tensor completion for better uncertainty quantification
A new research paper proposes a novel framework for evaluating large language models (LLMs) by treating pairwise human judgments as a tensor completion problem. This approach addresses the challenges of noisy, sparse, a…
-
New algorithm learns worker reliability and item rewards from comparisons
Researchers have developed a new algorithm to learn item rewards and worker reliability simultaneously from pairwise comparisons, particularly in crowdsourcing contexts. The proposed method utilizes a Boltzmann-rational…
-
Study finds global LLM leaderboards misleading, proposes portfolio rankings
A new research paper argues that current leaderboards for large language models (LLMs) are misleading due to significant heterogeneity in user preferences across languages and tasks. The study analyzed approximately 89,…