Researchers have developed a new three-stage framework to analyze how supervised fine-tuning (SFT), reinforcement learning (RL), and on-policy distillation (OPD) affect the confidence of large language models during reasoning. The study found that OPD is most effective for pre-reasoning confidence, SFT excels at early termination signals, and RL provides reliable trace-level confidence for answer aggregation. A novel position-aware confidence strategy, PosConf, was introduced to leverage confidence signals only from reliable relative-position intervals, improving answer aggregation and early stopping performance. AI
IMPACT Provides insights into improving LLM reliability and efficiency in reasoning tasks.
RANK_REASON Academic paper detailing a new analysis framework and strategy for LLM reasoning. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →