PulseAugur
EN
LIVE 07:58:15

NLI label variation study fails to replicate prior findings on monotonicity

A preregistered replication study aimed to verify a previously observed boundary in human label variation (HLV) within natural language inference (NLI) tasks. The original study suggested that hypotheses with non-upward monotonicity operators led to lower label agreement. However, this replication, conducted on unselected populations from the SNLI and MultiNLI development sets, failed to confirm this finding. Instead, the study observed slightly higher agreement for non-upward items, with all effect sizes falling below the threshold of interest, leading the authors to conclude that the original boundary might be conditional on selection methods rather than a general population property. AI

IMPACT This research highlights the importance of explicit reporting on data selection methods in NLI datasets, potentially influencing future dataset creation and evaluation practices.

RANK_REASON The cluster contains an academic paper detailing a scientific study and its findings. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

NLI label variation study fails to replicate prior findings on monotonicity

COVERAGE [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations

    Prior work on human label variation (HLV) in natural language inference (NLI) has often relied on re-annotation resources that select items by disagreement level. An earlier study (arXiv:2607.15870) found that hypotheses containing non-upward monotonicity operators showed lower l…