PulseAugur
EN
LIVE 21:48:23

MRT model advances layered image generation with 20B parameters

Researchers have introduced MRT, a 20-billion parameter masked region diffusion model designed for scalable layered image generation and editing. The model unifies text-to-layers, image-to-layers, and layers-to-layers tasks within a single framework, utilizing selective token masking for flexible layer manipulation. MRT also features an overflow-aware canvas to handle layers extending beyond visible boundaries and employs diffusion distillation for rapid, 8-step generation. AI

IMPACT Establishes a new benchmark for layered image generation, outperforming existing commercial systems and offering significant speed and memory improvements.

RANK_REASON The cluster describes a new research paper detailing a novel AI model for image generation.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

MRT model advances layered image generation with 20B parameters

COVERAGE [3]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

    A 20B-parameter masked region diffusion model enables scalable multi-layer transparent image generation and editing through unified task handling and efficient canvas management.

  2. arXiv cs.CV TIER_1 English(EN) · Zhicong Tang, Zhao Zhang, Jingye Chen, Mohan Zhou, Yifan Pu, Yuchi Liu, Yalong Bai, Ethan Smith, Yuhui Yuan ·

    MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

    arXiv:2605.27235v1 Announce Type: new Abstract: Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editing in natural language. Despite its importance, this …

  3. arXiv cs.CV TIER_1 English(EN) · Yuhui Yuan ·

    MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

    Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editing in natural language. Despite its importance, this remains an underexplored area at scale. To addre…