PulseAugur
实时 09:23:02

新框架结合人类判断和AI分数以改进评估

研究人员推出了一种名为Aggregate-then-Calibrate (AtC) 的新颖两阶段框架,旨在改进以人为中心的评估任务。该方法结合了异构的人类判断(考虑标注者可靠性)和模型生成的评分。AtC理论上证明,对标注者异构性的建模可以更有效地估计共识,并且其等渗校准即使在共识排名错误指定的情况下也能提供风险界限。实证结果表明,与仅依赖人类或模型输入的评估相比,AtC在准确性和鲁棒性方面始终表现更优。 AI

影响 该框架可以提高需要人类判断的领域中AI辅助决策过程的可靠性和准确性。

排序理由 该条目是一篇学术论文,详细介绍了一种用于评估任务的新框架。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv stat.ML 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新框架结合人类判断和AI分数以改进评估

报道来源 [1]

  1. arXiv stat.ML TIER_1 English(EN) · Zejun Xie, Xintong Li, Guang Wang, Desheng Zhang ·

    Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

    arXiv:2608.02455v1 Announce Type: cross Abstract: Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments…