arXiv:2407.07111v2 Announce Type: replace-cross Abstract: The rapid development of diffusion models (DMs) has significantly advanced image and video applications, making "what you want is what you see" a reality. Among these, video editing has gained substantial attention and see…
Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-fr…
Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-fr…
Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which …
Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus of the other annota…
arXiv:2610.03510v1 Announce Type: cross Abstract: Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backg…
arXiv:2311.18837v2 Announce Type: replace-cross Abstract: Diffusion models have achieved significant success in image and video generation. This motivates a growing interest in video editing tasks, where videos are edited according to provided text descriptions. However, most exi…
arXiv cs.AI
TIER_1English(EN)·Huawei Lin, Tony Geng, Zhaozhuo Xu, Weijie Zhao·
arXiv:2505.13439v2 Announce Type: replace-cross Abstract: Autoregressive (AR) models have recently shown strong performance in image generation, where a critical component is the visual tokenizer (VT) that maps continuous pixel inputs to discrete token sequences. The quality of t…
arXiv cs.AI
TIER_1English(EN)·Ziqi Ma, Shreya Sharma, Mohamed El Banani, Katja Schwarz, Chongjie Ye, Chao-Yuan Wu, Li Fei-Fei, Ben Mildenhall, Georgia Gkioxari, Justin Johnson, Gowthami Somepalli·
arXiv:2610.03636v1 Announce Type: cross Abstract: Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Exis…
Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which …
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 4…
Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while front-end decoding late…
arXiv cs.LG
TIER_1English(EN)·Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu·
arXiv:2609.35763v3 Announce Type: replace Abstract: Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce a unified theoretical framework that separates distribu…
Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects o…
Streaming video generation has benefited from distribution matching distillation (DMD), which matches the joint distribution of video frames to a video teacher's approximation of the real video distribution. Although this joint matching mitigates drift during autoregressive rollo…
arXiv cs.AI
TIER_1English(EN)·Amine Ouasfi, Runjia Li, Junlin Han, Eric Marchand, Philip H. S. Torr, Adnane Boukhayma·
arXiv:2609.39504v1 Announce Type: cross Abstract: We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffu…
arXiv:2609.37925v1 Announce Type: cross Abstract: Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. H…
arXiv cs.AI
TIER_1English(EN)·Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An, Haoyu Li, Yuhang Wu, Sen Cui, Miao Liu·
arXiv:2609.35491v2 Announce Type: replace-cross Abstract: Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution di…
Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing ap…
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow …
Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, c…
Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching…
Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every spe…
arXiv:2610.08684v1 Announce Type: new Abstract: Modern video models can generate realistic videos from real appearance references and proxy renders that specify scene structure, viewpoint changes, and motion. Evaluating this render-to-real capability requires a real target video …
arXiv cs.CV
TIER_1English(EN)·Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh, Xinyu Zhang·
arXiv:2610.08777v1 Announce Type: new Abstract: Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires …
arXiv:2610.08772v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive. Window attention offers an effi…
arXiv cs.CV
TIER_1English(EN)·Yunseung Ok (Kyung Hee University), Hyunsoo Kim (The University of Texas at Austin), Minseo Kim (Kyung Hee University), Suhyun Kim (Kyung Hee University)·
arXiv:2610.02914v1 Announce Type: new Abstract: Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-…
arXiv:2610.03221v1 Announce Type: new Abstract: Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distill…
arXiv:2610.03543v1 Announce Type: new Abstract: Streaming video generation has benefited from distribution matching distillation (DMD), which matches the joint distribution of video frames to a video teacher's approximation of the real video distribution. Although this joint matc…
arXiv:2610.02779v1 Announce Type: new Abstract: In this paper, we present trajectory-aware reuse and adaptive correction (TRAC), a training-free framework for efficient autoregressive (AR) video generation. Existing acceleration methods mainly target single-trajectory generation …
arXiv cs.CV
TIER_1English(EN)·Jeongwoo Shin, Youngyoon Choi, Sangwoo Jo, Hyunmog Kim, Sungjoon Choi, Joonseok Lee, Jaewoong Choi, Jaemoo Choi·
arXiv:2610.03120v1 Announce Type: new Abstract: Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works pr…
arXiv cs.CV
TIER_1English(EN)·Yiwen Zhang, Haocheng Xi, Michael Tian-Yue Liu, Alexei A. Efros, Hadar Averbuch-Elor, Qianqian Wang, Haiwen Feng·
arXiv:2610.02153v1 Announce Type: new Abstract: Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visual details may be lost and difficult to recover upon reappearance. To retain acce…
arXiv cs.CV
TIER_1English(EN)·Mikhail Dereviannykh, Vikram Voleti, Simon Donne, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss·
arXiv:2610.00686v1 Announce Type: new Abstract: Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in ter…
arXiv cs.CV
TIER_1English(EN)·Zhuo Ning, AmirHossein Naghi Razlighi, Sagi Polaczek, Daniel Cohen-Or, Ali Mahdavi-Amiri·
arXiv:2610.00691v1 Announce Type: new Abstract: Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workfl…
arXiv cs.CV
TIER_1English(EN)·Jiho Jang, Jinyoung Kim, Nojun Kwak, Kyungjune Kim·
arXiv:2610.01039v1 Announce Type: new Abstract: While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To…
arXiv cs.CV
TIER_1English(EN)·Tongcheng Zhang, Jun Zhu, Jianfei Chen·
arXiv:2610.01052v1 Announce Type: new Abstract: We argue that as video generation extends to longer durations, subject consistency should be evaluated over \textit{dynamic subject sets}. We therefore introduce \textbf{DynSC-Eval}, an evaluation framework that dynamically tracks e…
arXiv:2610.01661v1 Announce Type: new Abstract: Despite remarkable progress, video generation models often produce highly similar outputs when repeatedly sampled from the same prompt, limiting their usefulness for creative exploration. Existing diversity evaluations primarily rel…
arXiv cs.CV
TIER_1English(EN)·Cusuh Ham, Fabian Caba Heilbron, Josef Sivic, Bryan Russell·
arXiv:2610.01884v1 Announce Type: new Abstract: We introduce an approach for collection-grounded B-roll sequence generation. Given a user's video collection, a directive given in natural language, and a target duration, the goal is to produce a multi-shot sequence that complement…
arXiv:2610.02197v1 Announce Type: new Abstract: Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The pr…
arXiv cs.CV
TIER_1English(EN)·Yuxin Cao, Wei Song, Shangzhi Xu, Jingling Xue, Jin Song Dong·
arXiv:2602.10639v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what models predict, leaving the stability of how they generate largely unexamined. We …
arXiv:2609.38691v1 Announce Type: new Abstract: Streaming video generators allow users to dynamically modulate video synthesis via mid-stream prompt switching. Existing streaming methods can respond to the updated instruction while still cutting corners, prematurely realizing goa…
arXiv cs.CV
TIER_1English(EN)·Byoungwoo Park, Jaemoo Choi, Juho Lee, Yongxin Chen·
arXiv:2609.38562v1 Announce Type: new Abstract: World models, game simulators, and long-take video creation require coherent scene evolution and sustained dynamics over extended durations. Autoregressive (AR) video diffusion provides a natural framework for long-horizon generatio…
arXiv:2609.40037v1 Announce Type: new Abstract: Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context and can propagate through subsequent rollouts, leading to detail degradation, struc…
arXiv cs.CV
TIER_1English(EN)·Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng·
arXiv:2609.39132v1 Announce Type: new Abstract: We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a co…
arXiv:2609.38839v1 Announce Type: new Abstract: Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effec…
arXiv:2609.38154v1 Announce Type: new Abstract: Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation…
<table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1wxlyq2/pdmd_projected_distribution_matching_distillation/"> <img alt="PDMD: Projected Distribution Matching Distillation for Video Diffusion Models" src="https://external-preview.redd.it/OWZubGE1ZmFvaHRo…