Stanford Natural Language Inference corpus
PulseAugur coverage of Stanford Natural Language Inference corpus — every cluster mentioning Stanford Natural Language Inference corpus across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
New framework improves NLU model robustness via failure-mode bandits
Researchers have developed a novel adversarial data curation framework designed to enhance the robustness of natural language understanding models. This system frames data curation as a failure-mode contextual bandit pr…
-
New method generates commonsense axioms for NLI tasks, boosting LLM accuracy
Researchers have developed a method to generate commonsense knowledge axioms for Natural Language Inference (NLI) tasks, evaluating their effectiveness using LLMs like Llama 3.1 70B and GPT-OSS 120B. A novel reference-f…
-
NLI label variation study fails to replicate prior findings
A preregistered replication study on the Stanford Natural Language Inference corpus, MultiNLI, and ChaosNLI datasets has failed to confirm prior findings regarding human label variation. The original research suggested …
-
NLI label variation study fails to replicate prior findings on monotonicity
A preregistered replication study aimed to verify a previously observed boundary in human label variation (HLV) within natural language inference (NLI) tasks. The original study suggested that hypotheses with non-upward…
-
Formal semantic structure explains minimal human label variation in NLI tasks
A new research paper explores the extent to which formal semantic structure explains human label variation in natural language inference (NLI) tasks. The study analyzed items from the SNLI and MNLI corpora, finding that…
-
New research finds temperature scaling fails on soft labels
A new research paper challenges the effectiveness of temperature scaling for model calibration, particularly when dealing with soft or distributional human labels. The study found that temperature scaling, which assumes…
-
New audit protocol tests NLP benchmarks for evidence dependence
Researchers have developed a new auditing protocol for weak-label benchmarks in natural language processing. This protocol distinguishes between outputs predictable from metadata alone and those genuinely dependent on t…