Diffusion Transformers
PulseAugur coverage of Diffusion Transformers — every cluster mentioning Diffusion Transformers across labs, papers, and developer communities, ranked by signal.
9 day(s) with sentiment data
-
New BAG policy accelerates Diffusion Transformers with adaptive caching
Researchers have developed BAG (Budget-Aware Gating), a new caching policy designed to accelerate Diffusion Transformers (DiTs). Unlike previous methods that either lacked budget awareness or instance adaptivity, BAG us…
-
MAVISEG enhances zero-shot segmentation in diffusion transformers
Researchers have developed MAVISEG, a novel training-free refinement layer designed to enhance zero-shot open-vocabulary segmentation in diffusion transformers. Unlike existing methods that score pixels independently, M…
-
Flash-VAED framework accelerates video generation by 6x
Researchers have developed Flash-VAED, a framework designed to accelerate the VAE decoders used in latent diffusion models for video generation. This approach employs channel pruning and dominant operator optimization t…
-
New method detects physics plausibility in video diffusion models
Researchers have developed a method to identify physical plausibility in video diffusion models by analyzing intermediate denoising representations. They found that these models encode signals predictive of physical acc…
-
AdaCorrection framework boosts Diffusion Transformer efficiency for image generation
Researchers have developed AdaCorrection, a new framework designed to improve the efficiency of Diffusion Transformers (DiTs) used for image and video generation. DiTs are known for their high-quality output but are com…
-
New research explores gradient accuracy in downscaled image training for diffusion models
Researchers have investigated how training diffusion transformers on downscaled images affects gradient signals. They found that while a downscaled latent can preserve most of the surviving signal at high noise levels, …
-
New UDT Architecture Merges U-Nets and Diffusion Transformers
Researchers have introduced UDT, a novel architecture that merges the strengths of U-Nets and Diffusion Transformers (DiTs) for generative modeling. UDT employs data-adaptive token merging to reconcile the encoder-decod…
-
New DiverseDiT++ framework enhances Diffusion Transformer representation learning
Researchers have developed DiverseDiT++, a new framework designed to enhance the performance of Diffusion Transformers (DiTs) by promoting diversity in their internal representations. The study introduces a novel metric…
-
New ROAD framework slashes 3D generation training costs
Researchers have introduced ROAD, a novel framework designed to significantly reduce the computational costs associated with high-fidelity 3D shape generation. By leveraging the semantic and structural understanding fro…
-
New MIND network uses Diffusion Transformers for enhanced medical image fusion
Researchers have developed MIND, a novel Multimodal Intent-Driven Network that utilizes Diffusion Transformers for medical image fusion. This network integrates information from various imaging modalities by first using…
-
New methods enhance diffusion transformer efficiency and performance · 4 sources tracked
Researchers have developed new methods to improve the efficiency and performance of diffusion transformers, a key architecture for AI image and video generation. Chimera, a hybrid visual diffusion backbone, combines dif…
-
MegaSlide-DiT enables large video diffusion model adaptation on single GPU
Researchers have developed MegaSlide-DiT, a system enabling the adaptation of large video diffusion models on a single high-end GPU. This is achieved by keeping model weights and optimizer states in host RAM and streami…
-
Sol-Attn speeds up video generation with efficient sparse attention
Researchers have developed Sol-Attn, a new training-free sparse attention method designed to accelerate inference for video generation models. Unlike previous methods that struggle with efficiency and accuracy due to ri…
-
Diffusion Transformer study reveals hidden role of text template tokens
Researchers have developed a new interpretability framework for text-to-image diffusion transformers (DiTs) that reveals the crucial role of structural text template tokens. These tokens, often overlooked, act as implic…
-
New methods enhance multimodal control in Diffusion Transformers for image generation
Researchers have developed new methods to enhance control over image generation using Diffusion Transformers (DiTs). One approach, 'Appearance Pointers,' uses compact tokens to guide DiTs for precise regional control ov…
-
FlowSonic enables stable zero-shot music editing with diffusion transformers
Researchers have developed FlowSonic, a new framework for zero-shot music editing using diffusion transformers trained with rectified flow. This method aims to balance semantic changes with the preservation of a music r…
-
New PE-Field 4D framework enhances geometry control in video generation
Researchers have developed PE-Field 4D, a novel framework that enhances video generation models by improving control over scene geometry and viewpoint changes. This approach leverages positional encoding within Diffusio…
-
DiTango framework boosts Diffusion Transformer efficiency with selective attention
Researchers have developed DiTango, a new framework designed to make Diffusion Transformers (DiTs) more efficient for generating high-resolution content. DiTango addresses the scalability issues in parallelizing DiT inf…
-
VideoRAE leverages VFM features for improved video generation
Researchers have introduced VideoRAE, a novel representation autoencoder designed to enhance video generative models. This system leverages features from frozen Video Foundation Models (VFMs) like V-JEPA 2 and VideoMAEv…
-
VideoRAE enhances generative video models using frozen foundation features
Researchers have introduced VideoRAE, a novel representation autoencoder designed to enhance generative video modeling. Unlike traditional methods that focus on pixel-level reconstruction, VideoRAE leverages multi-scale…