English(EN)Discrete Diffusion Bridges for Spatiotemporally Aligned Image Translation and Generation
新研究探索用于图像和视频生成的扩散模型进展 · 跟踪9个来源
作者PulseAugur 编辑部·[9 个来源]·
arXiv上发布的多篇研究论文探讨了扩散模型在图像和视频生成方面的进展。这些研究引入了新颖的技术,例如用于矢量扩散图的基于地标的加速、用于图像增强的基于张量的泛函以及用于非线性扩散滤波的通道表示。其他论文侧重于提高扩散模型的效率和泛化能力,包括快速采样、弹性令牌压缩以及用于像素空间生成的航点扩散变换器等方法。此外,研究还通过候选回收和用于时空对齐图像翻译的离散扩散桥来解决视频扩散模型的测试时缩放问题。
AI
arXiv:2603.21247v2 Announce Type: replace-cross Abstract: We propose a landmark-constrained algorithm, LA-VDM (Landmark Accelerated Vector Diffusion Maps), to accelerate the Vector Diffusion Maps (VDM) framework built upon the Graph Connection Laplacian (GCL), which captures pair…
arXiv cs.CV
TIER_1English(EN)·Freddie {\AA}str\"om, Michael Felsberg, George Baravdish·
arXiv:2608.29164v1 Announce Type: new Abstract: In this work, we introduce a novel tensor-based functional for targeted image enhancement and denoising. Via explicit regularization, our formulation incorporates application dependent and contextual information using first principl…
arXiv cs.CV
TIER_1English(EN)·Christian Heinemann, Freddie {\AA}str\"om, George Baravdish, Kai Krajsek, Michael Felsberg, Hanno Scharr·
arXiv:2608.29227v1 Announce Type: new Abstract: In this work we propose a novel non-linear diffusion filtering approach for images based on their channel representation. To derive the diffusion update scheme we formulate a novel energy functional using a soft-histogram representa…
arXiv:2608.29233v1 Announce Type: new Abstract: We present the winning solution to the ACM Multimedia 2026 Grand Challenge on Single-Image Guided Multi-Angle Image Synthesis. It ranks first among 293 registered teams; 56 teams obtained at least one scored submission on the public…
arXiv cs.CV
TIER_1English(EN)·Eduard Zamfir, Christian Reisswig, Zongwei Wu, Yongqin Xian, Radu Timofte·
arXiv:2608.29281v1 Announce Type: new Abstract: Natural images concentrate their detail in a small fraction of the frame, yet diffusion models spend a full token on every patch, in every layer and at every timestep. The waste is largest in pixel-space models, with no autoencoder …
arXiv:2608.29322v1 Announce Type: new Abstract: Recent video diffusion models have achieved remarkable generation quality, but high-fidelity results still largely depend on closed-source systems or costly large-scale infrastructure. Test-time scaling (TTS) offers a training-free …
arXiv:2608.29997v1 Announce Type: new Abstract: We propose Discrete Diffusion Bridges (DDB), a novel framework designed to resolve the fundamental spatiotemporal misalignment of standard discrete diffusion in image translation and generation. By corrupting data into a pure mask s…
arXiv cs.CV
TIER_1English(EN)·Zhenyu Zhou, Defang Chen, Siwei Lyu, Chun Chen, Can Wang·
arXiv:2603.00763v2 Announce Type: replace Abstract: Text-to-image diffusion models have achieved unprecedented success but still struggle to produce high-quality results under limited sampling budgets. Existing training-free sampling acceleration methods are typically developed i…
arXiv:2603.15132v3 Announce Type: replace Abstract: While recent Flow Matching models avoid the reconstruction bottlenecks of latent autoencoders by operating directly in pixel space, the raw pixel manifold provides little explicit semantic organization, making target-specific tr…