A new arXiv paper reveals that the common technique of self-consistency, which involves averaging multiple model outputs, can actually decrease accuracy for smaller large language models (LLMs) on challenging science problems. For models like Qwen2.5-7B and Llama-3-8B, majority voting on chains of thought led to lower problem-specific accuracy compared to using a single inference pass. The research indicates that confidence scores do not reliably correlate with correctness for these models on difficult scientific reasoning tasks. AI
IMPACT Challenges the efficacy of a common LLM inference technique, potentially impacting how smaller models are deployed for complex reasoning tasks.
RANK_REASON Academic paper detailing a novel finding about LLM behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →