PulseAugur
EN
LIVE 08:03:54

NLI label variation study fails to replicate prior findings

A preregistered replication study on the Stanford Natural Language Inference corpus, MultiNLI, and ChaosNLI datasets has failed to confirm prior findings regarding human label variation. The original research suggested that hypotheses with non-upward monotonicity operators would show lower label agreement. However, this replication found the opposite, with non-upward items exhibiting slightly higher agreement, and all observed effects were below the threshold of interest. The researchers concluded that the previously identified negative boundary in label agreement may be an artifact of selection methods used in re-annotation resources rather than a genuine population-level property. AI

IMPACT Challenges assumptions about human label variation in NLI datasets, potentially impacting future dataset creation and model evaluation.

RANK_REASON Academic paper detailing a replication study of prior research findings. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

NLI label variation study fails to replicate prior findings

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Haram Choi ·

    Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations

    arXiv:2607.19231v1 Announce Type: new Abstract: Prior work on human label variation (HLV) in natural language inference (NLI) has often relied on re-annotation resources that select items by disagreement level. An earlier study (arXiv:2607.15870) found that hypotheses containing …