Researchers have identified a phenomenon called Dominant-vs-Dominated (DvD) imbalance in text-to-image diffusion models, where one concept token frequently suppresses others during multi-concept generation. A new benchmark, DominanceBench, has been developed to study this issue, examining its roots in both data and internal model mechanics. The study found that concepts learned from visually homogeneous training data are more prone to dominance, with dominant tokens capturing attention early in the denoising process and hindering the representation of competing concepts. This dominance appears to be distributed across multiple attention heads rather than being localized. AI
IMPACT Identifies a key failure mode in diffusion models that may hinder reliable multi-concept generation.
RANK_REASON Research paper detailing a new failure mode in diffusion models and introducing a benchmark to study it. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Diffusion Models
- DominanceBench
- Gotit.pub
- Hayeon Jeong
- Hugging Face
- IArxiv
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →