Researchers have introduced MIST, a new benchmark designed to evaluate how well language models can selectively trust external signals. The benchmark presents reasoning items under four conditions: clean, misleading, correct-context, and irrelevant-context. A new metric, SC2W, measures how often a misleading signal causes a correct answer to become wrong. The proposed SCOPE method uses Direct Preference Optimization (DPO) to train models on failures across all four conditions, significantly reducing susceptibility to misleading information while maintaining accuracy with trustworthy context. AI
IMPACT This research could lead to more reliable AI systems that can better discern trustworthy information from misleading signals.
RANK_REASON The cluster describes a new academic paper introducing a benchmark and method for evaluating language models.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →