Researchers have introduced Trajectory-Augmented Policy Optimization (TAPO), a novel method for self-distillation in large language models. Unlike traditional methods that implicitly align distributions, TAPO explicitly constructs corrective trajectories. These trajectories retain erroneous reasoning up to the point of failure, then incorporate natural-language diagnoses and corrected reasoning. Experiments on AIME 2024, AIME 2025, and HMMT 2025 demonstrate that TAPO improves both initial reasoning and error-correction effectiveness compared to GRPO. AI
IMPACT Enhances LLM reasoning capabilities by providing more targeted error correction during training.
RANK_REASON The cluster describes a new research paper detailing a novel method for improving large language model reasoning through self-distillation.
- AIME 2024
- AIME 2025
- Grpo
- HMMT 2025
- Tapolca
- Trajectory-Augmented Policy Optimization
- Kullback–Leibler divergence
- Self-distillation
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →