Researchers have developed a new method to recover correct answers from large language models (LLMs) that fail reasoning tasks. This technique addresses the issue of "expression failures," where models possess the underlying reasoning capability but struggle to output it correctly due to structural biases. By fitting just two parameters on unlabeled examples, the method can improve accuracy by 9-34 points for models like Qwen3.5, and also shows effectiveness on OLMo-2-1B and Llama-3.1-8B. AI
IMPACT This research suggests that many LLM reasoning failures may be due to output expression issues rather than a lack of capability, potentially altering how LLM performance is benchmarked and interpreted.
RANK_REASON The cluster contains an academic paper detailing a new method for evaluating LLM reasoning capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →