Researchers have developed the KuaiRP series of role-playing models, focusing on simplified prompt engineering, stable output quality, integrated domain knowledge, and efficient deployment. To address the challenge of injecting domain-specific knowledge without sacrificing general capabilities, they propose a multi-stage training pipeline. This includes a standardized character template and SFT data pipeline, followed by reinforcement learning with a rule-based reward function to prevent degradation. Finally, a novel self-distillation method called Two-stage On-Policy Distillation with Cumulative-Divergence Decay is used to recover general agent abilities. AI
IMPACT Introduces novel techniques for balancing domain-specific knowledge with general capabilities in role-playing models.
RANK_REASON This is a technical report detailing a new model series and training methodology. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Cumulative-Divergence Decay
- DagsHub
- Gotit.pub
- Hugging Face
- KuaiRP
- ScienceCast
- Two-stage On-Policy Distillation
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →