PulseAugur
EN
LIVE 22:54:22

New training method boosts LLM reasoning efficiency via confidence signals

Researchers have developed a novel self-supervised training method that improves the reasoning efficiency of large language models without explicitly optimizing for shorter outputs. By training models to predict their confidence in answers at intermediate reasoning steps, the models learn to generate more concise reasoning traces. This approach, applied to models like Gemma, Qwen, Nemotron, and GPT-OSS, reduced token generation by up to 25% on various benchmarks while maintaining accuracy. The findings suggest that metacognitive signals, such as confidence, can lead to efficient reasoning as a byproduct, rather than requiring direct optimization for length. AI

IMPACT This approach could lead to more cost-effective LLM inference by reducing token generation without sacrificing accuracy.

RANK_REASON Academic paper detailing a new training methodology for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New training method boosts LLM reasoning efficiency via confidence signals

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu, Genta Indra Winata, Anirban Das, Soheil Feizi, Nima Chitsazan ·

    Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

    arXiv:2609.31619v1 Announce Type: new Abstract: Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encour…