Diffusion Transformer
PulseAugur coverage of Diffusion Transformer — every cluster mentioning Diffusion Transformer across labs, papers, and developer communities, ranked by signal.
- instance of Krea.2 90%
- used by Krea.2 90%
- used by Krea 2 Turbo 90%
- used by Wan-Animate-2 90%
- instance of Krea 2 Turbo 90%
- developed by Krea 2 Turbo 90%
- developed by Wan-Animate-2 90%
- used by text-to-video generation 80%
- used by Influence Flower 70%
- instance of Influence Flower 70%
- used by diffusers 70%
- used by Bernini 70%
15 day(s) with sentiment data
-
Video DeltaNet enhances video diffusion model efficiency with hybrid attention
Researchers have developed Video DeltaNet (VDN), a novel approach to enhance the efficiency of video diffusion models. VDN addresses the computational bottleneck caused by attention mechanisms in processing long video s…
-
New JiT-DDT architecture trains text-to-image models 3.6x faster
Researchers have developed JiT-DDT, a novel architecture that significantly accelerates the training of text-to-image diffusion models. By unifying the compression and generation modules into a single model, JiT-DDT ach…
-
EMODY Flow generates emotion-aware full-body motion from audio
Researchers have developed EMODY Flow, a novel framework for generating full-body motion that synchronizes with speech and emotional cues. This system addresses a limitation in existing models where emotion conditioning…
-
SlotDiT uses object-centric slots for better video generation and robotics
Researchers have developed SlotDiT, a novel text-guided Diffusion Transformer that utilizes object-centric representations for improved video generation and robotic applications. Unlike previous models that relied on pi…
-
New neural network Multi4D maps material interfaces with 98.82% accuracy
Researchers have developed Multi4D, a novel neural network framework designed to analyze complex material interfaces using four-dimensional scanning transmission electron microscopy (4D-STEM). This system integrates a D…
-
UniLayDiff Unifies Content-Aware Layout Generation with Diffusion Transformer
Researchers have introduced UniLayDiff, a novel Unified Diffusion Transformer designed for content-aware layout generation. This model aims to unify various layout generation tasks, such as those conditioned by element …
-
Lumina-OmniLV framework unifies over 100 low-level vision tasks
Researchers have introduced Lumina-OmniLV, a unified multimodal framework designed for a wide array of low-level vision tasks. This framework, built on a Diffusion Transformer architecture, can handle over 100 sub-tasks…
-
Marigold V2: Open-source monocular depth estimation model released
A new open-source model called Marigold V2 has been released, capable of monocular depth estimation through a diffusion transformer. This model can be fine-tuned in under a week on a single consumer GPU, with its code, …
-
Diffusion Transformer enhances multimodal brain state decoding
Researchers have introduced CoMA-DiT, a novel bidirectional cross-modal Diffusion Transformer designed to enhance multimodal brain state decoding. This model leverages paired modalities as sources of mutual generative s…
-
AI generates synthetic plankton images to improve rare species classification
Researchers have developed a method to generate synthetic plankton imagery using a multimodal taxonomic conditioning approach. This technique addresses the issue of severely long-tailed datasets in automated plankton im…
-
Researcher trains 210M text-to-image DiT from scratch on single GPU
An individual trained a 210 million parameter text-to-image diffusion transformer from scratch, completing the process in 3.5 days on a single GPU. Key findings from this experiment include the observation that learned …
-
New framework enhances 3D geometry generation with multi-teacher distillation
Researchers have introduced Flow3D-OPD, a novel post-training framework designed to enhance 3D geometry generation models that utilize flow-matching diffusion Transformers. This two-stage approach incorporates multi-tea…
-
Marigold V2 advances monocular depth estimation using diffusion transformers
Researchers have developed Marigold V2, an advancement in monocular depth estimation that repurposes diffusion transformer (DiT) architectures. This new method achieves sharper and more detailed depth maps by employing …
-
AuK: Open-source model unifies speech generation and editing
Researchers have introduced AuK, an open-source foundational model designed for both speech generation and editing. This model integrates natural language instructions and audio context, utilizing a multimodal large lan…
-
RefDiT framework enhances image generation with local attribute guidance
Researchers have introduced RefDiT, a new framework designed to improve reference-based image generation. This method addresses limitations in current models that struggle with complex scenes containing multiple objects…
-
New ReaDiT Guidance framework enhances control in AI image and video generation
Researchers have introduced ReaDiT Guidance, a novel framework designed to enhance control over image and video generation using Diffusion Transformer (DiT) models. This method leverages internal feature representations…
-
SPARK method enhances frozen DiT models for image super-resolution
Researchers have developed SPARK, a novel method for enhancing image super-resolution using frozen Diffusion Transformer (DiT) models. SPARK focuses on modulating a small number of dominant internal channels, identified…
-
New EraseSAE framework enables precise concept removal in text-to-video models
Researchers have developed EraseSAE, a new framework designed to precisely remove specific concepts from text-to-video diffusion models. This method utilizes sparse autoencoders to isolate and erase unwanted semantics a…
-
Krea2T Enhancer adds phrase-level attention control for Stable Diffusion prompts
A new node called "Attention-Weighted Phrases" has been added to the Krea2T Enhancer, a tool for Stable Diffusion. This feature allows users to selectively increase or decrease the attention given to specific words or p…
-
New Text-Audiobox framework enables alignment-free voice dubbing and dialogue synthesis
Researchers have developed Alignment-Free Text-Audiobox (Text-AB), a novel framework for voice dubbing and dialogue synthesis. This system utilizes a Diffusion Transformer with a flow-matching objective and operates wit…