PulseAugur
EN
LIVE 08:23:36

New research reveals human annotators exhibit 'Averaging Bias' in AI faithfulness evaluations

A new paper published on arXiv, titled "Averaging Bias: Human Faithfulness Annotations are not Locally Faithful," investigates the reliability of human annotations for evaluating the faithfulness of text summarization models. The study found that human annotators may exhibit an "Averaging Bias," where they label summaries as faithful even if they contain factual errors, rather than strictly adhering to a rule that requires every sentence to be supported by the source document. This bias was identified by comparing global human labels with per-sentence judgments from LLM judges across several benchmarks, suggesting a need for improved annotation designs to ensure trustworthy evaluations. AI

IMPACT Highlights potential flaws in current AI evaluation methods, suggesting a need for more rigorous human annotation designs for faithfulness.

RANK_REASON The cluster contains a single academic paper published on arXiv discussing a novel concept related to AI evaluation. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New research reveals human annotators exhibit 'Averaging Bias' in AI faithfulness evaluations

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Huajian Zhang, Yiyang Feng, Jiawei Zhou ·

    Averaging Bias: Human Faithfulness Annotations are not Locally Faithful

    arXiv:2608.00205v1 Announce Type: new Abstract: Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence ma…