English(EN)MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer
新方法增强了扩散Transformer在图像生成中的多模态控制
作者PulseAugur 编辑部·[6 个来源]·
研究人员开发了新的方法来增强使用扩散Transformer(DiTs)的图像生成控制。一种方法,“Appearance Pointers”,使用紧凑型token来指导DiTs,实现对图像元素的精确区域控制,为本地化多模态引导提供了一种模态无关的接口。另一项开发,“DuSPiT”,引入了一种双分支架构,在像素空间扩散Transformer中将全局推理与局部外观建模分开,旨在获得更丰富的细节和更好的效率。对“Pixel-Space Diffusion Transformers”的回顾强调了它们通过在共享token空间中处理原始像素、文本和条件,在实现高保真度生成和统一多模态系统方面的潜力。
AI
Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-based methods, strong regional controllability, makeup fidelity, and identity preservation remain challenging. The reasons are …
arXiv:2607.19344v1 Announce Type: cross Abstract: Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text pro…
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can na…
arXiv cs.CV
TIER_1English(EN)·Ziyi Wang, Siming Zheng, Yang Yang, Shusong Xu, Hao Zhang, Bo Li, Changqing Zou, Peng-Tao Jiang·
arXiv:2607.20924v1 Announce Type: new Abstract: Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-based methods, strong regional controllability, makeup fidelity, and identity p…
arXiv:2607.18510v1 Announce Type: new Abstract: Diffusion Transformers achieve strong image generation performance, but most operate in compressed latent spaces. Pixel-space diffusion avoids this information loss, yet existing approaches map each raw image patch to a single token…
arXiv cs.CV
TIER_1English(EN)·Renye Yan, Jikang Cheng, You Wu, Ling Liang, Wei Peng, Athanasios V. Vasilakos, Qingyu Zhao, Yu Zhang, Ehsan Adeli, Kilian M. Pohl, Guoying Zhao·
arXiv:2607.17585v1 Announce Type: new Abstract: Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while separate represe…