Researchers have developed a novel self-supervised training method that improves the reasoning efficiency of large language models without explicitly optimizing for shorter outputs. By training models to predict their confidence in answers at intermediate reasoning steps, the models learn to generate more concise reasoning traces. This approach, applied to models like Gemma, Qwen, Nemotron, and GPT-OSS, reduced token generation by up to 25% on various benchmarks while maintaining accuracy. The findings suggest that metacognitive signals, such as confidence, can lead to efficient reasoning as a byproduct, rather than requiring direct optimization for length. AI
IMPACT This approach could lead to more cost-effective LLM inference by reducing token generation without sacrificing accuracy.
RANK_REASON Academic paper detailing a new training methodology for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →