PulseAugur
实时 07:25:26
English(EN) Towards demystifying the creativity of diffusion models

新研究解决扩散模型效率和应用问题 · 追踪8个来源

近期研究探索了扩散模型的进展,重点在于提高其效率和在各个领域的适用性。FlashDiff 引入了自适应区域执行和调度,以降低图像、视频和音频生成的服务延迟并提高吞吐量。视频扩散模型中存在“序列性差距”挑战,即性能会随着因果链的延长而下降,这表明需要改进序列计算。其他工作提出了使用检索增强扩散 Transformer 进行异步时间序列预测的 ReDiTT,并探索了无需重新训练即可提高扩散模型多样性的方差校正时间偏移。此外,Singularity Space 等新框架通过复平面奇点表示信号,以获得更好的结构稳定性,而 DUNE 提供了一种无需训练的精炼方法来减少扩散模型中的伪影。 AI

影响 这些进展旨在提高扩散模型在图像生成、视频预测和时间序列分析等各个领域的效率、多样性和适用性。

排序理由 多篇在 arXiv 上发表的研究论文,详细介绍了扩散模型的新方法和分析。

在 Google AI / Research 阅读 →

AI 生成摘要 · Google Gemini · 来自 100 个来源。 我们如何撰写摘要 →

新研究解决扩散模型效率和应用问题 · 追踪8个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇在 arXiv 上发表的研究论文,详细介绍了扩散模型的新方法和分析。
Source corroboration
100 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
55 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+35 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [100]

  1. Google AI / Research TIER_1 English(EN) ·

    迈向揭示扩散模型创造力的过程

    Algorithms & Theory

  2. Hugging Face Blog TIER_1 English(EN) ·

    将 Nunchaku 4位扩散推理引入 Diffusers

  3. arXiv cs.LG TIER_1 English(EN) · Jinshu Huang, Yiming Jiang, Chunlin Wu ·

    从分数学习到离散采样:扩散模型端到端泛化性分析

    arXiv:2607.23226v1 Announce Type: new Abstract: Despite the empirical success of score-based diffusion models, a complete theoretical understanding of how finite-sample learning, network parameterization, and numerical discretization jointly dictate generative quality remains und…

  4. arXiv cs.LG TIER_1 English(EN) · Arisrei Lim, Yossi Gandelsman ·

    学习扩散模型的采样参数

    arXiv:2607.23488v1 Announce Type: new Abstract: Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then h…

  5. arXiv cs.AI TIER_1 English(EN) · Luca Ambrogioni, Giulio Franzese, Alberto Foresti, Gabriel Raya, Bac Nguyen, Georgios Batzolis, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai, Yuki Mitsufuji ·

    从原子到熵:凸集下的扩散模型训练最优噪声分配

    arXiv:2607.20540v1 Announce Type: cross Abstract: How should a diffusion model decide which noise levels to train on, and how much? Despite the importance of this choice, current noise schedules are based largely on heuristics or empirical tuning. Here, we develop a general stati…

  6. arXiv cs.AI TIER_1 English(EN) · Du Yin, Estrid He, Juli\'an Jer\'onimo Ba\~nuelos, Yang Yang, Feng Hu, Yuchen Luo, Hao Xue, Stephan Sigg, Flora Salim ·

    StrideDiffusion:加速时间序列生成扩散模型

    arXiv:2607.20545v1 Announce Type: new Abstract: Diffusion models have become competitive generators for time series, but their practical use is limited by the large number of sequential denoising steps required at inference time. Existing fast samplers typically use fixed or gene…

  7. arXiv cs.AI TIER_1 English(EN) · Jingyuan Li, Xiaoyi Jiang, Yixuan Jiang, Wei Liu, Yi Zhu, Zuoqiang Shi, Pipi Hu ·

    均值到得分离散扩散:得分熵的后验均值去噪器

    arXiv:2607.21372v1 Announce Type: cross Abstract: Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity guarantees nonnegative reverse jump rates, it does not ensure Bayes realizability: ratios …

  8. arXiv cs.AI TIER_1 English(EN) · Yi Xiong, Yuan-Yuan Cheng, Xiao-Ming Fu ·

    面向高效扩散模型微调的源先验驱动选择性适应

    arXiv:2607.20913v1 Announce Type: new Abstract: Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained model's broad generative capability. Existing full and parameter-efficient fine-tu…

  9. arXiv cs.LG TIER_1 English(EN) · Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann ·

    KroQuant:用于扩散 Transformer 高效训练后量化的 Kronecker 结构块变换

    arXiv:2607.21446v1 Announce Type: new Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent. The standard fix applie…

  10. arXiv cs.AI TIER_1 English(EN) · Tianyi Zeng, Tianyi Wang, Jiaru Zhang, Zimo Zeng, Feiyang Zhang, Yiming Xu, Sikai Chen, Junfeng Jiao, Christian Claudel, Xinbo Chen ·

    PILD:基于扩散的物理信息学习

    arXiv:2601.21284v2 Announce Type: replace-cross Abstract: Diffusion models have emerged as powerful generative tools for modeling complex data distributions, yet their purely data-driven nature limits applicability in engineering and scientific problems where physical laws must b…

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    均值到得分离散扩散:得分熵的后验均值去噪器

    Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity guarantees nonnegative reverse jump rates, it does not ensure Bayes realizability: ratios at a noisy state need not be jointly induced by an…

  12. arXiv cs.LG TIER_1 English(EN) · Seonsoo Kim, Seongil Hong, Jun-Gill Kang ·

    Diffusion ReRoll:可修订去噪用于机器人序列预测

    arXiv:2607.19919v1 Announce Type: cross Abstract: We propose Diffusion ReRoll, a diffusion-based framework for robotic sequential prediction that enables revisable denoising over horizons. Existing diffusion-based sequence predictors typically perform a single monotonic denoising…

  13. arXiv cs.AI TIER_1 English(EN) · Kou Misaki, Takuya Akiba ·

    UnMaskFork: 通过确定性动作分支实现掩码扩散的测试时缩放

    arXiv:2602.04344v2 Announce Type: replace-cross Abstract: Test-time scaling strategies have effectively leveraged inference-time compute to enhance the reasoning abilities of Autoregressive Large Language Models. In this work, we demonstrate that Masked Diffusion Language Models …

  14. arXiv cs.AI TIER_1 English(EN) · Xiaoxuan Liang, Saeid Naderiparizi, Berend Zwartsenberg, Frank Wood ·

    集成很重要:约束扩散模型的基于推出的训练

    arXiv:2607.14398v1 Announce Type: cross Abstract: Constrained generative models aim to produce samples that satisfy complex feasibility constraints while remaining faithful to the data distribution. Existing constrained generation methods typically enforce constraints either thro…

  15. arXiv cs.LG TIER_1 English(EN) · Bernardo P. Schaeffer, Ricardo M. S. Rosa, Glauco Valle ·

    随机性在基于分数的扩散采样中的影响:KL散度分析

    arXiv:2506.11378v3 Announce Type: replace Abstract: Sampling in score-based diffusion models can be performed by solving either a reverse-time stochastic differential equation (SDE) parameterized by an arbitrary stochasticity function or a probability flow ODE, corresponding to s…

  16. arXiv cs.AI TIER_1 English(EN) · Parikshit Bansal, Sujay Sanghavi ·

    Token Time Continuous Diffusion for Language Modeling

    arXiv:2607.14106v1 Announce Type: cross Abstract: In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates in continuous space, deterministically mapping Gaussian noise to a final token canvas with no further sampling, a…

  17. arXiv cs.AI TIER_1 English(EN) · Ye Yuan, Weien Li, Rui Song, Zeyu Li, Haochen Liu, Xiangyu Kong, Zixuan Dong, Linfeng Du, Zipeng Sun, Weixu Zhang, Jiaxin Huang, Changjiang Han, Yonghan Yang, Zichen Zhao, Xiuyuan Hu, Haolun Wu, Yankai Chen, Fengran Mo, Jikun Kang, Bowei He, Philip S. Yu… ·

    离散扩散模型:从标记化到生成统一框架

    arXiv:2607.13431v1 Announce Type: cross Abstract: Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike cont…

  18. Hugging Face Daily Papers TIER_1 English(EN) ·

    离散扩散模型:从标记化到生成统一框架

    Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, …

  19. arXiv cs.AI TIER_1 English(EN) · Xue Liu ·

    离散扩散模型:从标记化到生成统一框架

    Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, …

  20. arXiv cs.LG TIER_1 English(EN) · Yaqi Qiao, Ping He, Songrun Xie, Ayush Barik, Chensong Zhang, Zhengzhong Tu, Fan Lai ·

    FlashDiff:高效区域执行与调度用于扩散模型服务

    arXiv:2607.12121v1 Announce Type: cross Abstract: Diffusion models have become the central backbone for modern image, video, and audio generation, but their efficient service remains a challenge. Unlike autoregressive decoding, diffusion inference repeatedly updates high-dimensio…

  21. arXiv cs.LG TIER_1 English(EN) · Saiyue Lyu, Zhitian Zhang, Ruizhi Deng, Thibaut Durand ·

    ReDiTT:用于异步时间序列的检索增强条件扩散 Transformer

    arXiv:2607.12391v1 Announce Type: new Abstract: We present a diffusion based model for asynchronous time series prediction, where the goal is to predict the next inter event time and event type. To address the inherent uncertainty of future events, we introduce ReDiTT, a retrieva…

  22. arXiv cs.LG TIER_1 English(EN) · Jorge Diaz Chao, Konpat Preechakul, Yuxi Liu, Yutong Bai ·

    视频扩散模型中的序列性鸿沟

    arXiv:2607.13031v1 Announce Type: new Abstract: When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video dif…

  23. Hugging Face Daily Papers TIER_1 English(EN) ·

    离散扩散模型:从标记化到生成统一框架

    Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, …

  24. Hugging Face Daily Papers TIER_1 English(EN) ·

    视频扩散模型中的序列性鸿沟

    When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, e…

  25. arXiv cs.LG TIER_1 English(EN) · Yutong Bai ·

    视频扩散模型中的序列性鸿沟

    When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, e…

  26. arXiv cs.LG TIER_1 English(EN) · Thibaut Durand ·

    ReDiTT:用于异步时间序列的检索增强条件扩散 Transformer

    We present a diffusion based model for asynchronous time series prediction, where the goal is to predict the next inter event time and event type. To address the inherent uncertainty of future events, we introduce ReDiTT, a retrieval augmented conditional diffusion transformer th…

  27. arXiv cs.LG TIER_1 English(EN) · Zichen Liu, Wei Zhang, Christof Sch\"utte, Tiejun Li ·

    黎曼去噪扩散概率模型

    arXiv:2505.04338v3 Announce Type: replace Abstract: We propose Riemannian Denoising Diffusion Probabilistic Models (RDDPMs) for learning distributions on submanifolds of Euclidean space that are level sets of functions, including most of the manifolds relevant to applications. Ex…

  28. arXiv cs.LG TIER_1 English(EN) · Peizhuo Li, Emre Aksan, Alexandru-Eugen Ichim, Thabo Beeler, Olga Sorkine-Hornung ·

    通过温度采样和方差校正时间偏移实现扩散模型多样化

    arXiv:2607.10853v1 Announce Type: cross Abstract: Diffusion models faithfully reproduce their training distribution, but also inherit its imbalances and leave rare or under-represented modes hard to reach. A natural inference-time remedy is to sample from the high-temperature tar…

  29. arXiv cs.AI TIER_1 English(EN) · Haksoo Lim, Myeongjin Lee, Wonjoon Chang, Jaesik Choi ·

    通过内部潜在分析对扩散模型进行统一骨干精炼

    arXiv:2607.09753v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success across diverse domains, with performance closely related to the denoising backbones that parameterize the score function. In this paper, we present a systematic, phase-aware analys…

  30. arXiv cs.AI TIER_1 English(EN) · Eli Bar-Yosef, Amir Averbuch, Eli Turkel ·

    奇点空间:用于信号表示的生成式扩散框架

    arXiv:2607.10930v1 Announce Type: cross Abstract: Generative models often represent signals as dense grids of amplitudes, blurring sharp transients that are crucial for the correctness of physical signals. We introduce Singularity Space, a generative framework that represents sig…

  31. arXiv cs.LG TIER_1 Italiano(IT) · Federico Ottomano, Gaopeng Ren, Yingzhen Li, Kim E. Jelfs, Alex M. Ganose ·

    用于3D分子生成的自回归潜在扩散模型

    arXiv:2607.09277v1 Announce Type: new Abstract: Three-dimensional (3D) molecule generation has been dominated by diffusion models, which achieve strong generation quality but typically require the molecular size to be specified a priori. Recent autoregressive approaches have subs…

  32. arXiv cs.AI TIER_1 English(EN) · Sang-Hoon Lee, Ha-Yeong Choi ·

    ReGen:高效波形扩散模型的分层多提示表示生成

    arXiv:2607.09134v1 Announce Type: cross Abstract: Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit genera…

  33. Hugging Face Daily Papers TIER_1 English(EN) ·

    Sticky Jump Diffusions: A Unifying View of Masked, Continuous, and Hybrid Diffusion

    We introduce Sticky Jump Diffusions (SJDs), continuous-time Markov processes on $\mathbb R^d$ whose discrete anchors are token embeddings. In forward time, anchors release their mass at a hazard rate and the released mass diffuses in the continuous ambient space; time reversal co…

  34. arXiv cs.LG TIER_1 Italiano(IT) · Alex M. Ganose ·

    用于3D分子生成的自回归潜在扩散模型

    Three-dimensional (3D) molecule generation has been dominated by diffusion models, which achieve strong generation quality but typically require the molecular size to be specified a priori. Recent autoregressive approaches have substantially narrowed the performance gap while nat…

  35. arXiv cs.AI TIER_1 English(EN) · Ha-Yeong Choi ·

    ReGen:高效波形扩散模型的分层多提示表示生成

    Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity. To address this issue, we propose R…

  36. arXiv cs.LG TIER_1 English(EN) · Abdullah Al Shafi, Sumaiya Rahim Suma ·

    弥合零空间:面向分类器自由扩散的引导感知量化

    arXiv:2607.08241v1 Announce Type: cross Abstract: Deploying classifier-free guidance (CFG) diffusion models under real-world compute budgets requires quantization, yet existing post-training quantization (PTQ) methods treat CFG models as single-branch networks, ignoring the paire…

  37. arXiv cs.LG TIER_1 English(EN) · Kaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang, Siyu Tang, Davis Rempe ·

    ARDY:用于交互式人类运动生成的混合表示自回归扩散模型

    arXiv:2607.08741v1 Announce Type: cross Abstract: Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinem…

  38. arXiv cs.LG TIER_1 English(EN) · Davis Rempe ·

    ARDY:用于交互式人类运动生成的混合表示自回归扩散模型

    Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While recent offline motion generation approaches offer precise control via text and kinematic constraints, they lack the inference speed re…

  39. Hugging Face Daily Papers TIER_1 English(EN) ·

    AutoAnchor:使用交叉注意力作为流形代理的Stable Diffusion去学习

    Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models. Current diffusion unlearning techniques determine the model update direction by either using alternatives of the target concept as an anchor or using empty p…

  40. arXiv cs.LG TIER_1 English(EN) · Sumaiya Rahim Suma ·

    弥合零空间:面向分类器自由扩散的引导感知量化

    Deploying classifier-free guidance (CFG) diffusion models under real-world compute budgets requires quantization, yet existing post-training quantization (PTQ) methods treat CFG models as single-branch networks, ignoring the paired conditional/unconditional structure that CFG inf…

  41. arXiv cs.LG TIER_1 English(EN) · Yi\u{g}it Berkay Uslu, Samar Hadou, Sergio Rozada, Shirin Saeedi Bidokhti, Alejandro Ribeiro ·

    随机图信号的生成式扩散模型

    arXiv:2607.06833v1 Announce Type: new Abstract: Sampling stochastic signals supported on a graph underlies many graph machine learning tasks, including recommender systems, forecasting in financial markets, and wireless network optimization. In these settings, the target signals …

  42. arXiv cs.AI TIER_1 English(EN) · Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay ·

    选择性时间步加权与基于优势的重放以实现样本高效的扩散式强化学习人类反馈

    arXiv:2607.07693v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly feedback inefficient, as existin…

  43. arXiv cs.AI TIER_1 English(EN) · Jinho Chang, Changsun Lee, Hyungjin Chung, Jong Chul Ye ·

    ContrastiveCFG: 通过对比正负概念引导扩散采样

    arXiv:2411.17077v2 Announce Type: replace-cross Abstract: As Classifier-Free Guidance (CFG) has proven effective in conditional diffusion model sampling for improved condition alignment, many applications use a negated CFG term as a Negative Prompting (NP) to filter out unwanted …

  44. arXiv cs.AI TIER_1 English(EN) · Zhiheng Zhou, Mengyao Zhou, Dengyi Zhao, Xingqin Qi, Guiying Yan ·

    Hypergraph Neural Stochastic Diffusion: An SDE Framework for Uncertainty Estimation

    arXiv:2607.07330v1 Announce Type: cross Abstract: Hypergraph neural networks have shown powerful capability in modeling higher-order relations, yet their predictive uncertainty remains underexplored. Unlike pairwise graphs, uncertainty in hypergraphs arises not only from noisy at…

  45. Hugging Face Daily Papers TIER_1 English(EN) ·

    增强多模态掩码扩散模型的生成顺序

    Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstrates that adaptive token generation ordering can significantly improve performance in mathematical reasoning and code synthesis applications.…

  46. Hugging Face Daily Papers TIER_1 English(EN) ·

    ARDY:用于交互式人类运动生成的混合表示自回归扩散模型

    ARDY is a streaming generation framework that enables real-time, high-fidelity 3D human motion generation with text and kinematic constraint control through a hybrid representation and two-stage autoregressive transformer denoiser.

  47. arXiv cs.AI TIER_1 English(EN) · Soumik Mukhopadhyay ·

    选择性时间步加权和基于优势的回放以实现样本高效的扩散式强化学习人类反馈

    Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly feedback inefficient, as existing approaches typically require large amounts of hu…

  48. arXiv cs.AI TIER_1 English(EN) · Guiying Yan ·

    Hypergraph Neural Stochastic Diffusion: An SDE Framework for Uncertainty Estimation

    Hypergraph neural networks have shown powerful capability in modeling higher-order relations, yet their predictive uncertainty remains underexplored. Unlike pairwise graphs, uncertainty in hypergraphs arises not only from noisy attributes and ambiguous labels, but also from varia…

  49. arXiv cs.LG TIER_1 English(EN) · Bowen Xue, Zihan Min, Xingyang Li, Zhekai Zhang, Haocheng Xi, Lvmin Zhang, Maneesh Agrawala, Jun-Yan Zhu, Song Han, Yujun Lin, Muyang Li ·

    FourTune:迈向扩散模型全4位高效训练后优化

    arXiv:2607.05711v1 Announce Type: new Abstract: Diffusion models have become a dominant paradigm for high-quality generative modeling, while post-training is essential for adapting them to diverse downstream applications. However, post-training of large diffusion models is still …

  50. arXiv cs.AI TIER_1 English(EN) · Wenhao Wang, Yifan Sun, Zongxin Yang, Zhengdong Hu, Zhentao Tan, Yi Yang ·

    视觉扩散模型中的复制:调查与展望

    arXiv:2408.00001v2 Announce Type: replace-cross Abstract: Visual diffusion models have revolutionized the field of creative AI, producing high-quality and diverse content. However, they inevitably memorize training images or videos, subsequently replicating their concepts, conten…

  51. Hugging Face Daily Papers TIER_1 English(EN) ·

    为基于扩散的模型运动生成检索和优化获胜噪声票据

    Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temporal, especially in compositional and long-duration sequences that require semantic consistency across multiple action segments …

  52. arXiv cs.LG TIER_1 English(EN) · Amandeep Kumar, Vishal M. Patel ·

    流形上的学习:用表示编码器解锁标准扩散 Transformer

    arXiv:2602.10099v2 Announce Type: replace Abstract: Leveraging representation encoders for generative modeling offers a path for efficient, high-fidelity synthesis. However, standard diffusion transformers fail to converge on these representations directly. While recent work attr…

  53. arXiv cs.AI TIER_1 English(EN) · Rajat Rasal, Avinash Kori, Tian Xia, Ben Glocker ·

    在扩散表示学习中引导优化轨迹

    arXiv:2607.05319v1 Announce Type: cross Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace this behaviour to optimisation dynamics; we analyse curves of image reconstruction against la…

  54. arXiv cs.LG TIER_1 English(EN) · Muyang Li ·

    FourTune:迈向扩散模型完全4位高效的训练后优化

    Diffusion models have become a dominant paradigm for high-quality generative modeling, while post-training is essential for adapting them to diverse downstream applications. However, post-training of large diffusion models is still challenging due to the prohibitive memory footpr…

  55. arXiv cs.AI TIER_1 English(EN) · Ben Glocker ·

    扩散表示学习中的优化轨迹引导

    We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace this behaviour to optimisation dynamics; we analyse curves of image reconstruction against latent representation quality, revealing trajectorie…

  56. Hugging Face Daily Papers TIER_1 English(EN) ·

    LILAC:用于扩散模型多概念定制的逐层独立LoRA和级联条件

    Personalizing text-to-image diffusion models to render several specific subjects in a coherent image remains challenging: the model must preserve each subject's identity while keeping the scene spatially and visually coherent. Methods that fuse independently trained concept adapt…

  57. Hugging Face Daily Papers TIER_1 English(EN) ·

    扩散模型和流匹配采样器的渐近保持后验分析

    Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $σ_{\min}$, at which the score is stiff and the flow develops a boundary layer. We treat $σ_{\min}$ as a singular-perturbation parameter and deter…

  58. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Xi Liu ·

    Diffusion-GR2:Diffusion生成式推理重排器

    Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trac…

  59. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Xi Liu ·

    Diffusion-GR2: Diffusion 生成式推理重排器

    Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autoregressive (AR) decoder spends one sequential forward pass per reasoning token, and the reasoning trac…

  60. arXiv stat.ML TIER_1 English(EN) · Lan V. Truong ·

    从分数近似到分布近似:基于分数的扩散模型

    arXiv:2607.22199v1 Announce Type: cross Abstract: Score-based diffusion models have achieved remarkable empirical success in generative modeling, yet their approximation-theoretic foundations remain incomplete. In particular, although classical universal approximation theorems gu…

  61. arXiv cs.CV TIER_1 English(EN) · Yifan Zhou, Zeqi Xiao, Tianyi Wei, Shuai Yang, Xingang Pan ·

    可训练对数线性稀疏注意力机制用于高效扩散Transformer

    arXiv:2512.16615v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) set the state of the art in visual generation, yet their quadratic self-attention cost fundamentally limits scaling to long token sequences. Recent Top-K sparse attention approaches reduce the compu…

  62. arXiv cs.CV TIER_1 English(EN) · Yidong Luo, Chenggong Li, Yuchao Feng, Boxin Shi, Junchao Zhang, Xin Yuan ·

    Stokes信息扩散用于鲁棒线性偏振估计

    arXiv:2607.21239v1 Announce Type: new Abstract: Polarization cues benefit applications such as material detection and de-reflection, yet acquiring them typically requires dedicated hardware. This motivates us to estimate the linear polarization from a single RGB image. However, t…

  63. arXiv cs.CV TIER_1 English(EN) · Rogerio Guimaraes, Pietro Perona ·

    通过渐进式种子剪枝实现扩散模型的推理时间缩放

    arXiv:2607.21591v1 Announce Type: new Abstract: Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for autoregressive language models. Because final quality is highly sensitive to the in…

  64. arXiv stat.ML TIER_1 English(EN) · Louis Grenioux, Maxence Noble ·

    基于扩散的退火玻尔兹曼生成器:益处、陷阱与希望

    arXiv:2601.21026v2 Announce Type: replace Abstract: Sampling configurations at thermodynamic equilibrium is a central challenge in statistical physics. Boltzmann Generators (BGs) tackle it by combining a generative model with a Monte Carlo (MC) correction step to obtain asymptoti…

  65. arXiv cs.CV TIER_1 English(EN) · Ba-Thinh Lam, Srijan Das, Hieu Le ·

    面向扩散模型的注意力感知OBS剪枝

    arXiv:2607.20048v1 Announce Type: new Abstract: We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps -- derived f…

  66. arXiv cs.CV TIER_1 English(EN) · Vaibhav Vavilala, Rahul Vasanth, David Forsyth ·

    使用扩散模型对蒙特卡洛渲染进行去噪

    arXiv:2404.00491v3 Announce Type: replace Abstract: Physically-based renderings contain Monte Carlo noise, with variance that increases as the number of rays per pixel decreases. This noise, while zero-mean for good modern renderers, can have heavy tails (most notably, for scenes…

  67. arXiv cs.CV TIER_1 English(EN) · Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng ·

    视频扩散模型能否预测过去帧?用于可逆插值的双向循环一致性

    arXiv:2604.01700v2 Announce Type: replace Abstract: Video frame interpolation aims to synthesize realistic intermediate frames between given endpoints while adhering to specific motion semantics. While recent generative models have improved visual fidelity, they predominantly ope…

  68. arXiv cs.CV TIER_1 English(EN) · Yumeng Ren, Yaofang Liu, Aitor Artola, Laurent Mertz, Raymond H. Chan, Jean-michel Morel ·

    通过截断的 Karhunen--Lo\`eve 展开改进扩散生成模型

    arXiv:2503.17657v3 Announce Type: replace Abstract: Pretrained diffusion models exhibit a well-known training-sampling mismatch, often attributed to exposure bias and related distribution-shift effects. We provide a quantitative interpretation of this phenomenon through the notio…

  69. arXiv cs.CV TIER_1 English(EN) · Weilai Xiang, Hongyu Yang, Di Huang, Yunhong Wang ·

    Conditioning Residuals for Diffusion Models via Representation Feedback

    arXiv:2505.10999v4 Announce Type: replace Abstract: Diffusion models now serve as a common foundation for multimedia generation, and useful intermediate representations emerge during their generative training. Standard architectures, however, propagate these representations throu…

  70. arXiv stat.ML TIER_1 English(EN) · Angus Phillips, Thomas Seror, Michael Hutchinson, Valentin De Bortoli, Arnaud Doucet, Emile Mathieu ·

    光谱扩散过程

    arXiv:2209.14125v3 Announce Type: replace Abstract: Diffusion models have proven to be a flexible and effective framework for modelling probability distributions on finite-dimensional spaces. However, many physical modelling problems such as time series are naturally described ov…

  71. arXiv stat.ML TIER_1 English(EN) · Han Chen, Sifan Liu, Jun Yang ·

    马尔可夫链蒙特卡洛与扩散路径

    arXiv:2607.11631v1 Announce Type: cross Abstract: Sampling from multimodal distributions is a longstanding challenge for classical local Markov chain Monte Carlo (MCMC) methods. A popular remedy is to introduce a sequence of intermediate distributions that interpolate between the…

  72. arXiv stat.ML TIER_1 English(EN) · Ziv Aharoni, Henry D. Pfister ·

    Diffusion Models 的守恒律

    arXiv:2607.10067v1 Announce Type: cross Abstract: While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typically trained with denoising objectives. We develop conservation laws based on generalized extrinsic information transfer …

  73. arXiv stat.ML TIER_1 English(EN) · Pascal Jutras-Dub\'e, Patrick Pynadath, Jeremy Lu, Yuan Gao, Ruqi Zhang ·

    Sticky Jump Diffusions: A Unifying View of Masked, Continuous, and Hybrid Diffusion

    arXiv:2607.10951v1 Announce Type: cross Abstract: We introduce Sticky Jump Diffusions (SJDs), continuous-time Markov processes on $\mathbb R^d$ whose discrete anchors are token embeddings. In forward time, anchors release their mass at a hazard rate and the released mass diffuses…

  74. arXiv stat.ML TIER_1 English(EN) · Lei Qian, Wu Su, Yanqi Huang, Song Xi Chen ·

    Likelihood Matching for Diffusion Models

    arXiv:2508.03636v3 Announce Type: replace Abstract: We propose a Likelihood Matching approach for training diffusion models by first establishing an equivalence between the likelihood of the target data distribution and a likelihood along the sample path of the reverse diffusion.…

  75. arXiv stat.ML TIER_1 English(EN) · Jiadong Liang, Zhihan Huang, Yuxin Chen ·

    扩散模型的低维适应:全变分收敛

    arXiv:2501.12982v3 Announce Type: replace Abstract: This paper investigates how diffusion generative models leverage (unknown) low-dimensional structure to accelerate sampling. Focusing on two mainstream samplers -- the denoising diffusion implicit model (DDIM) and the denoising …

  76. arXiv stat.ML TIER_1 English(EN) · Patrick Pynadath, Jiaxin Shi, Ruqi Zhang ·

    CANDI: 混合离散-连续扩散模型

    arXiv:2510.22510v3 Announce Type: replace-cross Abstract: While continuous diffusion has shown remarkable success in continuous domains such as image generation, its direct application to discrete data has underperformed pure discrete formulations. To understand this gap, we intr…

  77. arXiv cs.CV TIER_1 English(EN) · Yongseong Park, Joeun Kim, HoEun Kim, Young-Sik Kim ·

    噪声锚定扩散反演中的压缩不对称性和轨迹绑定

    arXiv:2607.09784v1 Announce Type: new Abstract: Real-image diffusion inversion is governed by a tight quality-cost trade-off, with costs incurred in computation, storage, or per-image optimization. We study this trade-off through the forward Gaussian noise anchor that defines a d…

  78. arXiv stat.ML TIER_1 English(EN) · Jun Yang ·

    马尔可夫链蒙特卡洛与扩散路径

    Sampling from multimodal distributions is a longstanding challenge for classical local Markov chain Monte Carlo (MCMC) methods. A popular remedy is to introduce a sequence of intermediate distributions that interpolate between the target and a simpler reference. The classical cho…

  79. arXiv cs.CV TIER_1 English(EN) · Yasong Dai, Zeeshan Hayder, David Ahmedt-Aristizabal, Hongdong Li ·

    探索扩散去噪动力学以进行对比表示学习

    arXiv:2607.09067v1 Announce Type: new Abstract: Text-to-image diffusion models exhibit unprecedented generative capability and contain rich intermediate representations that can be useful for discriminative vision tasks. Motivated by this observation, we study a focused question:…

  80. arXiv cs.CV TIER_1 English(EN) · Lingchen Sun, Rongyuan Wu, Zhengqiang Zhang, Ruibin Li, Yujing Sun, Shuaizheng Liu, Lei Zhang ·

    自我超越:外部特征引导对加速扩散Transformer训练是否不可或缺?

    arXiv:2601.07773v3 Announce Type: replace Abstract: Recent works such as REPA have shown that guiding diffusion models with external semantic features (e.g., DINO) can significantly accelerate the training of diffusion transformers (DiTs). However, the use of pretrained external …

  81. arXiv stat.ML TIER_1 English(EN) · Ruqi Zhang ·

    Sticky Jump Diffusions: A Unifying View of Masked, Continuous, and Hybrid Diffusion

    We introduce Sticky Jump Diffusions (SJDs), continuous-time Markov processes on $\mathbb R^d$ whose discrete anchors are token embeddings. In forward time, anchors release their mass at a hazard rate and the released mass diffuses in the continuous ambient space; time reversal co…

  82. arXiv stat.ML TIER_1 English(EN) · Henry D. Pfister ·

    Diffusion Models 的守恒律

    While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typically trained with denoising objectives. We develop conservation laws based on generalized extrinsic information transfer (GEXIT) functions for a broad class of memoryless …

  83. arXiv stat.ML TIER_1 English(EN) · Yidong Ouyang, Zhe Wang, Sourav Bhabesh, Dmitriy Bespalov ·

    多模态掩码扩散模型生成顺序的增强

    arXiv:2607.08056v1 Announce Type: cross Abstract: Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstrates that adaptive token generation ordering can significantly improve performance in mathe…

  84. arXiv cs.CV TIER_1 English(EN) · Cheng Wan, Bahram Jafrasteh, Ehsan Adeli, Miaomiao Zhang, Qingyu Zhao ·

    解剖学引导的潜在扩散模型用于脑部MRI进展建模

    arXiv:2601.14584v2 Announce Type: replace Abstract: Accurately modeling longitudinal brain MRI progression is crucial for understanding neurodegenerative diseases and predicting individualized structural changes. Existing state-of-the-art approaches, such as Brain Latent Progress…

  85. arXiv stat.ML TIER_1 English(EN) · Siyuan Wen, Jiahao Zeng, Ningning Ding ·

    AutoAnchor:使用交叉注意力作为流形代理的Stable Diffusion去学习

    arXiv:2607.08337v1 Announce Type: cross Abstract: Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models. Current diffusion unlearning techniques determine the model update direction by either using alternatives o…

  86. arXiv cs.CV TIER_1 English(EN) · Hongdong Li ·

    探索扩散去噪动力学以进行对比表示学习

    Text-to-image diffusion models exhibit unprecedented generative capability and contain rich intermediate representations that can be useful for discriminative vision tasks. Motivated by this observation, we study a focused question: how can the denoising dynamics of a pretrained …

  87. arXiv stat.ML TIER_1 English(EN) · Ningning Ding ·

    AutoAnchor:使用交叉注意力作为流形代理的稳定扩散解学

    Diffusion unlearning is essential for mitigating the generation of harmful or copyrighted content in text-to-image models. Current diffusion unlearning techniques determine the model update direction by either using alternatives of the target concept as an anchor or using empty p…

  88. arXiv stat.ML TIER_1 English(EN) · Robert Gruhlke, Julius Berner, David Sommer, Lorenz Richter ·

    Tensor Train Diffusion:利用低秩结构进行高维基于分数的采样

    arXiv:2607.06841v1 Announce Type: new Abstract: Diffusion models offer a powerful framework for sampling from complex probability densities by learning to reverse a noising process. A common approach involves solving for the time-reversed stochastic differential equation (SDE), w…

  89. arXiv stat.ML TIER_1 English(EN) · Dmitriy Bespalov ·

    多模态掩码扩散模型生成顺序的增强

    Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstrates that adaptive token generation ordering can significantly improve performance in mathematical reasoning and code synthesis applications.…

  90. arXiv cs.CV TIER_1 English(EN) · Wanglong Lu, Lingming Su, Kaijie Shi, Minglun Gong, Xiaogang Jin, Hanli Zhao, Xianta Jiang ·

    用于超高分辨率图像编辑的无微调潜在扩散模型

    arXiv:2607.06136v1 Announce Type: new Abstract: Recent diffusion-based generative models have shown impressive performance in image generation and editing. However, due to memory limitations and the high cost of collecting high-resolution training images, existing methods are typ…

  91. arXiv stat.ML TIER_1 English(EN) · Lorenz Richter ·

    Tensor Train Diffusion:利用低秩结构进行高维基于分数的采样

    Diffusion models offer a powerful framework for sampling from complex probability densities by learning to reverse a noising process. A common approach involves solving for the time-reversed stochastic differential equation (SDE), which requires the score function of the evolving…

  92. arXiv cs.CV TIER_1 English(EN) · Xianta Jiang ·

    用于超高分辨率图像编辑的无微调潜在扩散模型

    Recent diffusion-based generative models have shown impressive performance in image generation and editing. However, due to memory limitations and the high cost of collecting high-resolution training images, existing methods are typically restricted to inputs with linear resoluti…

  93. arXiv cs.CV TIER_1 English(EN) · Chunnan Shang, Xin Zhang, Zhizhong Wang, Hongwei Wang ·

    DICT:用于扩散模型条件图像生成的注入和对比轨迹细化数据

    arXiv:2607.03899v1 Announce Type: new Abstract: Diffusion models have become a dominant paradigm for conditional image generation, yet existing approaches generally follow two directions: task-specific designs that can improve performance but limit generalization, and training-fr…

  94. arXiv cs.CV TIER_1 English(EN) · Aryan Das, Koushik Biswas, Moloud Abdar, Vinay Kumar Verma ·

    UNITY:用于扩散模型自适应条件化的注意力流网络

    arXiv:2606.20971v2 Announce Type: replace Abstract: We introduce UNITY, a Universal-to-Specialized adapter for efficient and scalable composite conditioning in diffusion based image generation. Unlike prior methods that train separate adapters for each conditioning modality, UNIT…

  95. arXiv cs.CV TIER_1 English(EN) · Lijiang Li, Zuwei Long, Yunhang Shen, Heting Gao, Haoyu Cao, Xing Sun, Caifeng Shan, Ran He, Chaoyou Fu ·

    Omni-Diffusion:基于掩码离散扩散的统一多模态理解与生成

    arXiv:2603.06577v2 Announce Type: replace Abstract: While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective and effici…

  96. Sequoia Capital TIER_1 English(EN) · sbarry ·

    与Sable合作:缩小扩散差距

    <p>The post <a href="https://sequoiacap.com/article/partnering-with-sable-closing-the-diffusion-gap/">Partnering with Sable: Closing the Diffusion Gap</a> appeared first on <a href="https://sequoiacap.com">Sequoia Capital</a>.</p>

  97. dev.to — LLM tag TIER_1 Русский(RU) · Promptra Team ·

    stable diffusion 9009 及其他数字:为什么不存在“SD4”以及该用什么替代

    <p>Чёрное окно консоли, строка <code>Couldn't launch python</code>, финал: <code>exit code: 9009</code>. Stable Diffusion не стартует, а поисковик по запросу «stable diffusion 9009» выдаёт всё подряд - от форумов по Windows до «новостей» про выход SD4. Спойлер сразу: 9009 - систе…

  98. dev.to — LLM tag TIER_1 English(EN) · Gemini Team ·

    DiffusionGemma: 开发者指南

    <p>Following our announcement in our <a href="https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/" rel="noopener noreferrer">launch blog post</a>, we are sharing this developer guide to help you understand, serve and customize…

  99. r/MachineLearning TIER_1 English(EN) · /u/Savings-Display5123 ·

    LingBot-Video:稀疏MoE视频扩散Transformer(总计13B,活跃1.4B)作为动作条件世界模型进行后训练[R]

    <!-- SC_OFF --><div class="md"><p>Single-stream diffusion transformer with a DeepSeek-V3-style sparse MoE (128 experts, top-8 routing, 1.4B active of 13B total). Six-reward RL post-training including a physical-plausibility reward, plus an action-to-video mode that predicts robot…

  100. r/StableDiffusion TIER_2 (AF) · /u/Course_Latter ·

    在hfviewer中可视化扩散模型

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1uss13d/visualizing_diffusion_models_in_hfviewer/"> <img alt="Visualizing diffusion models in hfviewer" src="https://external-preview.redd.it/YnU4cWE0am5lZmNoMcXtDsnCthSSb3yXDEZEpJhpeywVzsasg5ljAh9UCiuV.p…