New research explores advanced diffusion models for generation, robustness, and speed
ByPulseAugur Editorial·[51 sources]·
Researchers are developing advanced diffusion models for various applications, including image generation, time-series synthesis, and natural language processing. New methods like Simplax aim to improve categorical generation by augmenting states with auxiliary variables, while PhysDGM embeds physical laws into diffusion models for realistic time-series data synthesis in dynamic systems. Other advancements focus on enhancing adversarial robustness in diffusion models through techniques like gradient masking and space compression, and accelerating inference speeds for diffusion transformers using novel forecasting methods. Additionally, new frameworks are being explored for faster sampling of diffusion models and for developing diffusion large language models that balance accuracy with parallelism.
AI
IMPACT
These advancements in diffusion models could lead to more realistic data generation, improved robustness against adversarial attacks, and faster inference for complex AI tasks.
RANK_REASON
Multiple research papers published on arXiv detailing new methods and frameworks for diffusion models.
arXiv:2505.16733v3 Announce Type: replace Abstract: This paper proposes to perform image restoration through a state-dependent mean-reverting forward diffusion (FoD) process. In contrast to traditional diffusion-based approaches that rely on a coupled forward-backward diffusion s…
arXiv cs.CL
TIER_1English(EN)·Jinya Sakurai, Patrick Pynadath, Satoshi Hayakawa, Jaehong Yoon, Xulei Yang, Nancy F. Chen, Xun Xu·
arXiv:2608.10615v1 Announce Type: new Abstract: Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem. We study uniform discrete diffusion and ask whethe…
arXiv:2505.22839v2 Announce Type: replace-cross Abstract: Recent studies suggest that diffusion models significantly improve the empirical adversarial robustness of deep neural network models. While intuitive explanations have been proposed, the mechanisms underlying diffusion-ba…
arXiv cs.AI
TIER_1English(EN)·Hu Yu, Hao Luo, Xueyang Fu, Jie Huang, Fan Wang, Feng Zhao·
arXiv:2506.13058v2 Announce Type: replace-cross Abstract: Diffusion probabilistic models (DPMs) have demonstrated remarkable success in visual generation. However, their iterative sampling mechanism results in slow inference speeds. While reducing sampling steps offers an intuiti…
arXiv cs.LG
TIER_1English(EN)·Haiteng Wang, Yunfei Zhu, Tao Wang, Yikang Li, Jiabao Dong, Xiaoge Zhang, Lei Ren·
arXiv:2608.10941v1 Announce Type: new Abstract: Industrial time-series signals, such as turbine temperature and rotational speed in aero-engines, are essential for monitoring the health and operational status of complex dynamical systems. However, collecting such data is often li…
arXiv:2608.09133v1 Announce Type: cross Abstract: Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging. In particular, we observe that diffusion transformer…
arXiv cs.AI
TIER_1Deutsch(DE)·Shiyi Qi, Kun He, Mingmou Liu·
arXiv:2608.08594v1 Announce Type: new Abstract: Diffusion bridge models leverage Doob's \(h\)-transform to construct stochastic transports between arbitrary endpoint distributions, and have shown strong potential in image-to-image translation and restoration. However, most existi…
arXiv:2603.27996v2 Announce Type: replace Abstract: Diffusion models have emerged as a powerful framework for generative tasks in deep learning. They decompose generative modeling into two computational primitives: deterministic neural-network evaluation and stochastic sampling. …
arXiv cs.AI
TIER_1English(EN)·Jinlong Yang, Jinke Wu, Lizilin, Yao Zhou·
arXiv:2608.07572v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive computational overhead, temporal feature caching has been proposed to bypass redundant…
Simplax enriches uniform discrete diffusion via Dirichlet-categorical augmentation to improve reverse sampling and generative quality on text and Sudoku tasks.
arXiv:2608.07161v1 Announce Type: cross Abstract: Simulating complex fluid flows requires capturing full equilibrium distributions rather than just mean trajectories, yet high-fidelity solvers remain computationally prohibitive. Recent advances, such as Diffusion Graph Networks (…
arXiv:2601.07568v3 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) offer capabilities beyond those of autoregressive (AR) LLMs, such as parallel decoding and random-order generation. However, realizing these benefits in practice is non-trivial, as d…
Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences. Reinforcement learning (RL) for preference alignment in diffusi…
arXiv cs.LG
TIER_1English(EN)·Wenhan Guo, Jinglun Yu, Yaning Wang, Jin U. Kang, Yu Sun·
arXiv:2512.18367v2 Announce Type: replace-cross Abstract: Diffusion models are highly expressive image priors for Bayesian inverse problems. However, most diffusion models cannot operate on large-scale, high-dimensional data due to high training and inference costs. In this work,…
arXiv cs.LG
TIER_1English(EN)·Ruchi Sandilya, Sumaira Perez, Charles Lynch, Lindsay Victoria, Benjamin Zebley, Derrick Matthew Buchanan, Mahendra T. Bhati, Nolan Williams, Timothy J. Spellman, Faith M. Gunning, Conor Liston, Logan Grosenick·
arXiv:2510.14190v3 Announce Type: replace Abstract: Diffusion models excel at generation, but their latent spaces are high dimensional and not explicitly organized for interpretation or control. We introduce ConDA (Contrastive Diffusion Alignment), a plug-and-play geometry layer …
We introduce the Intrinsic Hybrid Latent Diffusion Model (ILDM), a generative framework that integrates probabilistic dimensionality reduction with geometry-aware diffusion on unknown manifolds. While diffusion models (DMs) have achieved state-of-the-art results in high-dimension…
arXiv:2608.03117v1 Announce Type: new Abstract: The performance of generative diffusion models is determined by the choice of the reference diffusion process connecting the empirical and prior distributions. Conventional approaches typically trade off simulation-free training aga…
Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, …
Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then held fixed across prompts and denoising timesteps…
arXiv:2505.19196v2 Announce Type: replace Abstract: Recent advances in text-to-image (T2I) diffusion model fine-tuning leverage reinforcement learning (RL) to align generated images with learnable reward functions. The existing approaches reformulate denoising as a Markov decisio…
arXiv:2608.13205v1 Announce Type: new Abstract: Text-Image-to-Video (TI2V) models are an emerging unified architecture, where a single model simultaneously supports text-to-video (T2V) and image-to-video (I2V) generation. Given a high-quality first frame or a detailed textual pro…
arXiv cs.CV
TIER_1English(EN)·Xilong Zhou, Pedro Figueiredo, Milo\v{s} Ha\v{s}an, Valentin Deschaintre, Paul Guerrero, Yiwei Hu, Nima Khademi Kalantari·
arXiv:2509.01134v2 Announce Type: replace-cross Abstract: Generative models for high-quality materials are particularly desirable to make 3D content authoring more accessible. However, the majority of material generation methods are trained on synthetic data. Synthetic data provi…
arXiv:2303.08063v4 Announce Type: replace Abstract: For a considerable time, researchers have focused on developing a method that establishes a deep connection between the generative diffusion model and mathematical physics. Despite previous efforts, progress has been limited to …
arXiv stat.ML
TIER_1English(EN)·Martin J. Wainwright·
arXiv:2608.13520v1 Announce Type: cross Abstract: We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \emph{unmasking growth complexity} ({\textsf{UGC}\xspace}). Its local increments directly control Kullback--Leibler…
arXiv:2602.23783v5 Announce Type: replace Abstract: Text-to-image (T2I) diffusion models lack an efficient mechanism for early quality assessment, leading to costly trial-and-error in multi-generation scenarios such as prompt iteration, agent-based generation, and flow-grpo. We r…
arXiv:2608.12155v1 Announce Type: new Abstract: Vision foundation models have recently emerged as powerful feature extractors for detecting AI-generated images, achieving strong generalization across generators and robustness to common image degradations. However, the reason behi…
arXiv cs.CV
TIER_1English(EN)·Shizhuo Mao, Hongtao Zou, Qihu Xie, Song Chen, Yi Kang·
arXiv:2512.05746v3 Announce Type: replace Abstract: Diffusion models have demonstrated significant applications in the field of image generation. However, their high computational and memory costs pose challenges for deployment. Model quantization has emerged as a promising solut…
arXiv:2608.09445v1 Announce Type: cross Abstract: Unconditional diffusion checkpoint merging assumes benign sources, yet a compromised public checkpoint can transfer a dormant backdoor while clean generation appears normal. Mitigation is difficult without knowing the compromised …
arXiv:2608.08770v1 Announce Type: new Abstract: Diffusion models are increasingly used as controllable samplers, whose generations can be steered at inference time according to a chosen reward function. While such rewards are typically defined on individual samples, for many appl…
arXiv:2608.08519v1 Announce Type: new Abstract: Intensity-image reconstruction from event streams remains a challenging problem due to the binary, sparse, and asynchronous nature of event data. This work proposes eBIRD, an event-guided reconstruction framework that combines a DDP…
arXiv:2602.18093v2 Announce Type: replace Abstract: Diffusion Transformers (DiT) have emerged as a widely adopted backbone for high-fidelity image and video generation, yet their iterative denoising process incurs high computational costs. Existing training-free acceleration meth…
arXiv cs.CV
TIER_1English(EN)·Yu Xue, Haoxuan Qu, Zhuoling Li, Hongbin Xu, Jianxiong Yin, Simon See, Hossein Rahmani, Jun Liu·
arXiv:2608.07003v1 Announce Type: new Abstract: Training-free text-to-high-resolution image generation has recently attracted growing research attention. However, existing studies on this task primarily focus on adapting off-the-shelf U-Net-based diffusion models to high resoluti…
arXiv:2605.07327v2 Announce Type: replace Abstract: Sampling from pretrained diffusion and flow-matching models typically requires many forward passes to generate diverse and high-fidelity images. Existing distillation methods often rely on multiple auxiliary networks, carefully …
arXiv:2608.06768v1 Announce Type: new Abstract: Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences. Reinforcement l…
arXiv:2510.22778v3 Announce Type: replace-cross Abstract: We develop a free-probabilistic framework for denoising diffusion, in which the data is a self-adjoint operator and its law a spectral distribution. The forward process is the free Ornstein--Uhlenbeck diffusion, whose spec…
arXiv:2608.06794v1 Announce Type: new Abstract: While diffusion models have made significant progress in text-to-image tasks, they still exhibit limitations when directly optimizing downstream objectives. Although Reinforcement Learning (RL) enables targeted optimization, existin…
arXiv stat.ML
TIER_1English(EN)·Jairon H. N. Batista, Fl\'avio B. Gon\c{c}alves, Yuri F. Saporito, Rodrigo S. Targino·
arXiv:2505.06800v2 Announce Type: replace Abstract: Diffusion-based generative models have renewed interest in stochastic differential equation methods for sampling from complex distributions. We study a setting in which the target density is known only up to a normalizing consta…
arXiv cs.CV
TIER_1English(EN)·Rui Li, Yuanzhi Liang, Ke Hao, Ziqiao Weng, Haibin Huang, Chi Zhang, XueLong Li·
arXiv:2608.06125v1 Announce Type: new Abstract: Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. However, existing latent reward models output only scalar …
arXiv stat.ML
TIER_1English(EN)·Yizhu Wang, Mu Niu, Xiaochen Yang·
arXiv:2608.04827v1 Announce Type: new Abstract: We introduce the Intrinsic Hybrid Latent Diffusion Model (ILDM), a generative framework that integrates probabilistic dimensionality reduction with geometry-aware diffusion on unknown manifolds. While diffusion models (DMs) have ach…
arXiv stat.ML
TIER_1Italiano(IT)·Mahsa Taheri, Johannes Lederer·
arXiv:2502.09151v3 Announce Type: replace-cross Abstract: Diffusion models are one of the key architectures of generative AI. Their main drawback, however, is the computational costs. This study indicates that the concept of sparsity, well known especially in statistics, can prov…
arXiv stat.ML
TIER_1English(EN)·Yufei Wu, Shanqing Gao, Andreas Voss, Francis Tuerlinckx·
arXiv:2608.03566v1 Announce Type: new Abstract: The drift diffusion model (DDM) is a cornerstone of cognitive decision-making research. Although numerous estimation methods exist, researchers continue to seek inference approaches that are both fast and flexible across diverse stu…
arXiv cs.CV
TIER_1English(EN)·Seokho Han, Dongwei Wang, Jinhee Kim, Yiran Chen, Kang Eun Jeon, Huanrui Yang, Jong Hwan Ko·
arXiv:2608.03057v1 Announce Type: new Abstract: Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sensitive step, even though many other steps can tolerate fewer bits. The resulting…
arXiv:2502.06606v3 Announce Type: replace Abstract: Manipulating the material appearance of objects in images is critical for applications like augmented reality, virtual prototyping, and digital content creation. We present MaterialFusion, a novel framework for high-quality mate…
arXiv:2608.03469v1 Announce Type: new Abstract: We study score learning for reflected diffusion on bounded domains. Reflection keeps trajectories feasible but does not ensure that the learned score satisfies the boundary behavior implied by the forward process. With implicit scor…
arXiv:2507.21449v2 Announce Type: replace Abstract: Degeneracy is an inherent feature of the loss landscape of neural networks, but it is not well understood how stochastic gradient MCMC (SGMCMC) algorithms interact with this degeneracy. In particular, existing global convergence…
arXiv:2512.08022v2 Announce Type: replace Abstract: We propose a novel diffusion-based posterior sampling method within a plug-and-play framework. Our approach constructs a probability transport from an easy-to-sample distribution to the target posterior via a diffusion process. …
arXiv cs.CV
TIER_1English(EN)·Ankit Yadav, Ta Duc Huy, Lingqiao Liu·
arXiv:2512.17303v2 Announce Type: replace Abstract: In diffusion and flow-matching generative models, guidance techniques are widely used to improve sample quality and consistency. Classifier-free guidance (CFG) is the de facto choice in modern systems and achieves this by contra…
<p>После утверждения иллюстрация может остаться только результатом: выбранный визуал есть, а общей записи исходных материалов, ограничений и последующих изменений нет. В работе со stable diffusion нейросетью это превращает просьбу сделать ещё один вариант в риск случайно заменить…