A preregistered replication study aimed to verify a previously observed boundary in human label variation (HLV) within natural language inference (NLI) tasks. The original study suggested that hypotheses with non-upward monotonicity operators led to lower label agreement. However, this replication, conducted on unselected populations from the SNLI and MultiNLI development sets, failed to confirm this finding. Instead, the study observed slightly higher agreement for non-upward items, with all effect sizes falling below the threshold of interest, leading the authors to conclude that the original boundary might be conditional on selection methods rather than a general population property. AI
IMPACT This research highlights the importance of explicit reporting on data selection methods in NLI datasets, potentially influencing future dataset creation and evaluation practices.
RANK_REASON The cluster contains an academic paper detailing a scientific study and its findings. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- arXiv:2607.15870
- ChaosNLI
- Hugging Face Daily Papers
- MultiNLI
- Selection Shapes the Boundary: A Preregistered Replication of Monotonicity and Label Agreement in Unselected NLI Populations
- Stanford Natural Language Inference corpus
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →