Researchers have developed a new post-training method called Reward-Weighted Transport Distillation (RWTD) for one-step generative models. This technique allows for alignment using only generated samples and scalar reward evaluations, addressing the difficulty in training these models. RWTD constructs an adaptive target that balances reward adaptation with the retention of prior knowledge by mixing separately tilted current and reference distributions. Empirically, RWTD improved the GenEval score of the SANA Sprint 1.6B backbone from 0.73 to 0.80 and demonstrated strong generalization across different rewards. AI
IMPACT This new method could improve the training efficiency and performance of one-step generative models, potentially leading to better image and media generation.
RANK_REASON The cluster contains an academic paper detailing a new method for generative models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →