Researchers have introduced ToPO (Token-Oriented Preference Optimization), a novel method for training latent diffusion models using pairwise image preferences. Unlike previous methods that apply preferences to complete images, ToPO constructs a spatial-temporal route to incorporate these preferences during the denoising process. This approach utilizes content tokens to modulate cross-attention and includes an auxiliary pixel-midpoint ordering term, eliminating the need for local labels or a learned reward model. Evaluations show ToPO outperforms Diffusion-DPO on multiple metrics for both SD-1.5 and SDXL models, demonstrating higher endpoint estimates and larger win shares in comparative studies. AI
IMPACT This new method could lead to more efficient and effective training of generative image models by better leveraging pairwise preferences.
RANK_REASON The cluster contains a research paper detailing a new method for training diffusion models. [lever_c_demoted from research: ic=1 ai=1.0]
- Diffusion-DPO
- HPSv2
- ImageReward
- SD Negeri 1.5 Belimbing
- SDXL
- Token-Oriented Preference Optimization
- ToPO
- U-Net
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →