A new paper explores the effectiveness of On-Policy Distillation (OPD) for enhancing large language models, particularly focusing on data efficiency and selection. The research found that even a single example (1-shot OPD) can be effective, with harder examples often yielding superior performance gains. The study suggests that improvements stem from longer Chain-of-Thought (CoT) paths in complex problems, which help maintain alignment with the teacher model and teach critical thinking patterns. Based on these findings, a data selection method was proposed that uses only hard examples, achieving performance comparable to a much larger dataset with just 8 selected hard examples. AI
IMPACT Suggests a more data-efficient approach to training LLMs, potentially reducing computational costs and improving reasoning capabilities.
RANK_REASON Academic paper detailing a new method for training LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →