MMDiT
PulseAugur coverage of MMDiT — every cluster mentioning MMDiT across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Regional Prompter compatibility issues emerge with new AI image generation architectures
The Regional Prompter tool, a popular method for controlling image generation regions in Stable Diffusion, is facing compatibility issues with newer transformer architectures like Flux and DiT. While the core plugin rem…
-
Krea 2 Identity Edit: Unofficial model for instruction-based image editing released
A community fine-tune of the Krea 2 model, named Krea 2 Identity Edit, has been released on Hugging Face. This unofficial model allows for instruction-based image editing while preserving the identity of subjects, inclu…
-
New research explores diffusion models, bias mitigation, and reinforcement learning applications · 10 sources tracked
Recent research explores advancements in diffusion models, focusing on theoretical underpinnings, optimization techniques, and bias mitigation. One paper introduces Bayesian Information Restricted Diffusion (BIRD) model…
-
New AI models UniGP and AccelAes unify image generation and perception tasks
Researchers have developed UniGP, a framework that unifies controllable image generation and dense prediction tasks by jointly training a diffusion transformer model. This approach, built on MMDiT, aims to capture the j…
-
Hybrid Diffusion Transformer Enhances Instruction-Guided Audio Editing
Researchers have developed a novel hybrid diffusion transformer architecture for instruction-guided audio editing. This two-stage approach, based on rectified flow matching, aims to improve both the accuracy and efficie…
-
Alibaba's Qwen-RobotWorld Unifies Embodied AI with Language Interface
Alibaba's Qwen team has introduced Qwen-RobotWorld, a language-conditioned video world model designed for embodied intelligence. This model utilizes natural language as a universal interface to predict future visual tra…
-
New frameworks and benchmarks advance audio-visual generation
Researchers have introduced OmniCustom, a framework for customizing both video identity and audio timbre simultaneously from reference images and audio. This DiT-based model uses separate LoRA modules for identity and t…
-
New framework creates lightweight diffusion models via knowledge distillation
Researchers have developed a new knowledge distillation framework called LIFT and PLACE to create more efficient diffusion models. This method addresses the difficulty students have in mimicking complex teacher models b…
-
New benchmarks challenge MLLMs' spatial and functional reasoning abilities
Researchers have introduced new benchmarks to evaluate the spatial and functional reasoning capabilities of multimodal large language models (MLLMs). These benchmarks aim to move beyond basic geometric perception to ass…
-
AttnRouter enhances image editing on MMDiT with per-category attention routing
Researchers have developed AttnRouter, a novel method for training-free image editing on the MMDiT model. This approach utilizes KVInject, a single-forward attention manipulation that blends source-image key/value proje…
-
OccDirector: Language-Guided Behavior and Interaction Generation in 4D Occupancy Space
Researchers have introduced OccDirector, a new framework designed to generate complex 4D occupancy dynamics for autonomous driving simulations based solely on natural language instructions. This system acts as a "scenar…