PulseAugur
中
实时 20:02:23
English(EN) MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer

新方法增强了扩散Transformer在图像生成中的多模态控制

研究人员开发了新的方法来增强使用扩散Transformer(DiTs)的图像生成控制。一种方法,“Appearance Pointers”,使用紧凑型token来指导DiTs,实现对图像元素的精确区域控制,为本地化多模态引导提供了一种模态无关的接口。另一项开发,“DuSPiT”,引入了一种双分支架构,在像素空间扩散Transformer中将全局推理与局部外观建模分开,旨在获得更丰富的细节和更好的效率。对“Pixel-Space Diffusion Transformers”的回顾强调了它们通过在共享token空间中处理原始像素、文本和条件,在实现高保真度生成和统一多模态系统方面的潜力。 AI

影响 这些在扩散Transformer方面的进展可能为创意专业人士和多模态应用带来更精确、更高效的AI驱动图像生成。

排序理由 该集群包含详细介绍扩散Transformer新架构和方法的学术论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

新方法增强了扩散Transformer在图像生成中的多模态控制

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含详细介绍扩散Transformer新架构和方法的学术论文。
Source corroboration
6 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
70 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [6]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    MagicMakeup:一种区域可控的扩散Transformer,用于高保真度美妆迁移

    Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-based methods, strong regional controllability, makeup fidelity, and identity preservation remain challenging. The reasons are …

  2. arXiv cs.AI TIER_1 English(EN) · Rahul Sajnani, Yulia Gryaditskaya, Radom\'ir M\v{e}ch, Srinath Sridhar, Matheus Gadelha ·

    外观指针 -- Diffusion Transformers的多模态区域控制

    arXiv:2607.19344v1 Announce Type: cross Abstract: Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text pro…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    外观指针 -- Diffusion Transformers的多模态区域控制

    Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can na…

  4. arXiv cs.CV TIER_1 English(EN) · Ziyi Wang, Siming Zheng, Yang Yang, Shusong Xu, Hao Zhang, Bo Li, Changqing Zou, Peng-Tao Jiang ·

    MagicMakeup:一种区域可控的扩散Transformer,用于高保真度美妆迁移

    arXiv:2607.20924v1 Announce Type: new Abstract: Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-based methods, strong regional controllability, makeup fidelity, and identity p…

  5. arXiv cs.CV TIER_1 English(EN) · Yunpeng Bai, Yossi Gandelsman, Micha\"el Gharbi ·

    DuSPiT: 双分支子块像素扩散 Transformer

    arXiv:2607.18510v1 Announce Type: new Abstract: Diffusion Transformers achieve strong image generation performance, but most operate in compressed latent spaces. Pixel-space diffusion avoids this information loss, yet existing approaches map each raw image patch to a single token…

  6. arXiv cs.CV TIER_1 English(EN) · Renye Yan, Jikang Cheng, You Wu, Ling Liang, Wei Peng, Athanasios V. Vasilakos, Qingyu Zhao, Yu Zhang, Ehsan Adeli, Kilian M. Pohl, Guoying Zhao ·

    Pixel-Space Diffusion Transformers

    arXiv:2607.17585v1 Announce Type: new Abstract: Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while separate represe…