PulseAugur
中
实时 16:06:52
English(EN) Soundwich: Video Generation with Layered and Controllable Audio

新研究推动可控且对齐的视频生成 · 跟踪10个来源

视频生成领域的最新研究正致力于提高控制和对齐能力,以应对时间连贯性和符合人类意图等挑战。几篇论文介绍了新的训练后和对齐方法,包括监督微调、自训练和基于奖励的技术。这些进展旨在提高生成视频的质量和可控性,其中一些方法专注于特定方面,如摄像机控制、音频生成或交互合成。 AI

影响 可控且对齐的视频生成方面的进展可能会加速媒体、机器人和VR/AR领域的应用。

排序理由 多篇arXiv论文详细介绍了视频生成的新方法和调查。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 49 个来源。 我们如何撰写摘要 →

新研究推动可控且对齐的视频生成 · 跟踪10个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇arXiv论文详细介绍了视频生成的新方法和调查。
Source corroboration
49 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
8 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+25 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [49]

  1. arXiv cs.AI TIER_1 English(EN) · Wenhao Sun, Rong-Cheng Tu, Jingyi Liao, Dacheng Tao ·

    基于扩散模型的视频编辑:一项调查

    arXiv:2407.07111v2 Announce Type: replace-cross Abstract: The rapid development of diffusion models (DMs) has significantly advanced image and video applications, making "what you want is what you see" a reality. Among these, video editing has gained substantial attention and see…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    CtrlCache:通过感知控制的缓存加速交互式视频世界模型

    Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-fr…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    CtrlCache:通过感知控制的缓存加速交互式视频世界模型

    Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-fr…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    S2PD:用于物理和逻辑一致视频生成的串行到并行扩散

    Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which …

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    TasteRoute:视频生成的个性化路由

    Rapid progress in video generation has led to a plethora of models that differ substantially in capability and generation cost. This raises a natural question: can each request be efficiently routed to an appropriate model? We find that even when the consensus of the other annota…

  6. arXiv cs.AI TIER_1 English(EN) · Ziyi Wang, Junchi Yao, Heqian Qiu, Wenbo Shi, Chengjiu Wang, Jinyang He, Binkai Hong, Hongliang Li ·

    Weave Forcing:用于交互式长视频生成的组合式记忆路由

    arXiv:2610.03510v1 Announce Type: cross Abstract: Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backg…

  7. arXiv cs.AI TIER_1 English(EN) · Zhen Xing, Shuyuan Tu, Qi Dai, Zihao Zhang, Hui Zhang, Han Hu, Zuxuan Wu, Yu-Gang Jiang ·

    VIDiff:通过多模态指令和扩散模型进行视频翻译

    arXiv:2311.18837v2 Announce Type: replace-cross Abstract: Diffusion models have achieved significant success in image and video generation. This motivates a growing interest in video editing tasks, where videos are edited according to provided text descriptions. However, most exi…

  8. arXiv cs.AI TIER_1 English(EN) · Huawei Lin, Tony Geng, Zhaozhuo Xu, Weijie Zhao ·

    VTBench:评估自回归图像生成的视觉分词器

    arXiv:2505.13439v2 Announce Type: replace-cross Abstract: Autoregressive (AR) models have recently shown strong performance in image generation, where a critical component is the visual tokenizer (VT) that maps continuous pixel inputs to discrete token sequences. The quality of t…

  9. arXiv cs.AI TIER_1 English(EN) · Ziqi Ma, Shreya Sharma, Mohamed El Banani, Katja Schwarz, Chongjie Ye, Chao-Yuan Wu, Li Fei-Fei, Ben Mildenhall, Georgia Gkioxari, Justin Johnson, Gowthami Somepalli ·

    LoGo:用于一致性长时域视频生成的局部-全局奖励

    arXiv:2610.03636v1 Announce Type: cross Abstract: Camera-controlled video models are rapidly advancing toward long generation horizons and complex camera control. A key failure mode is 3D inconsistency: as the camera moves, objects lose permanence and scene structures shift. Exis…

  10. Hugging Face Daily Papers TIER_1 English(EN) ·

    S2PD:用于物理和逻辑一致视频生成的串行到并行扩散

    Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which …

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    Kandinsky 6.0 视频:用于同步视频和音频生成的基石模型

    We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 4…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    重新思考长视频效率:帧、像素和前端延迟的联合分配视角

    Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while front-end decoding late…

  13. arXiv cs.LG TIER_1 English(EN) · Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu ·

    统一分布训练以实现一步视觉生成

    arXiv:2609.35763v3 Announce Type: replace Abstract: Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce a unified theoretical framework that separates distribu…

  14. arXiv cs.AI TIER_1 English(EN) · Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli ·

    视频生成模型:训练后和对齐的调查

    arXiv:2610.00812v1 Announce Type: cross Abstract: Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pret…

  15. Hugging Face Daily Papers TIER_1 English(EN) ·

    测试时用于长视频生成的分布内强制

    Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects o…

  16. Hugging Face Daily Papers TIER_1 English(EN) ·

    DuoMatching:用于少样本视频生成的联合边缘分布匹配

    Streaming video generation has benefited from distribution matching distillation (DMD), which matches the joint distribution of video frames to a video teacher's approximation of the real video distribution. Although this joint matching mitigates drift during autoregressive rollo…

  17. arXiv cs.AI TIER_1 English(EN) · Amine Ouasfi, Runjia Li, Junlin Han, Eric Marchand, Philip H. S. Torr, Adnane Boukhayma ·

    PartiCam:通过奖励引导进行相机控制的视频生成

    arXiv:2609.39504v1 Announce Type: cross Abstract: We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffu…

  18. arXiv cs.AI TIER_1 English(EN) · Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang, Weidong Zhang, Tianfan Xue ·

    面向长视域自回归视频生成的Rollout-Marginal蒸馏

    arXiv:2609.37925v1 Announce Type: cross Abstract: Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. H…

  19. arXiv cs.AI TIER_1 English(EN) · Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An, Haoyu Li, Yuhang Wu, Sen Cui, Miao Liu ·

    从分数到样本:自回归视频生成的弹性强制

    arXiv:2609.35491v2 Announce Type: replace-cross Abstract: Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution di…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    FrameMorrow:面向长视频生成的未来引导帧选择与前瞻性Token

    Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing ap…

  21. Hugging Face Daily Papers TIER_1 English(EN) ·

    视频生成模型:训练后和对齐的调查

    Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow …

  22. Hugging Face Daily Papers TIER_1 English(EN) ·

    SemanTok:高效自回归视频生成的可预测语义令牌

    Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, c…

  23. Hugging Face Daily Papers TIER_1 English(EN) ·

    面向长视域自回归视频生成的Rollout-Marginal蒸馏

    Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching…

  24. Hugging Face Daily Papers TIER_1 English(EN) ·

    LongLive-Plug:面向视频生成的一次性蒸馏

    Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every spe…

  25. arXiv cs.CV TIER_1 English(EN) · Dicong Qiu, Zhiyuan Xu, Yaosheng Liu, Feng Han, Bo Ye ·

    RenderBench:使用重建的数字孪生对渲染到真实视频传输进行基准测试

    arXiv:2610.08684v1 Announce Type: new Abstract: Modern video models can generate realistic videos from real appearance references and proxy renders that specify scene structure, viewpoint changes, and motion. Evaluating this render-to-real capability requires a real target video …

  26. arXiv cs.CV TIER_1 English(EN) · Shangye Song, Dong Gong, Hong Jia, Yun Sing Koh, Xinyu Zhang ·

    CtrlCache:通过感知控制的缓存加速交互式视频世界模型

    arXiv:2610.08777v1 Announce Type: new Abstract: Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires …

  27. arXiv cs.CV TIER_1 English(EN) · Liao Ma, Jiayi Song, Yunfeng Wu, Songhua Liu, Peilin Zhao ·

    面向快速高分辨率视觉生成的后端无关稀疏注意力机制

    arXiv:2610.08772v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have achieved strong performance in image and video generation, but the quadratic complexity of full attention makes high-resolution generation computationally expensive. Window attention offers an effi…

  28. arXiv cs.CV TIER_1 English(EN) · Yunseung Ok (Kyung Hee University), Hyunsoo Kim (The University of Texas at Austin), Minseo Kim (Kyung Hee University), Suhyun Kim (Kyung Hee University) ·

    定制强制:自回归视频生成无需训练的主题定制

    arXiv:2610.02914v1 Announce Type: new Abstract: Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-…

  29. arXiv cs.CV TIER_1 English(EN) · Yutong Wang, Xingtong Ge, Enhuai Liu, Yunke Wang, Tianfan Xue, Yu Qiao, Yaohui Wang, Xinyuan Chen, Chang Xu ·

    VDOT++:通过不平衡最优传输蒸馏实现统一的少样本视频生成

    arXiv:2610.03221v1 Announce Type: new Abstract: Video creation spans text-to-video (T2V), image-to-video (I2V), and condition-based generation, yet video diffusion models remain costly because they repeatedly evaluate large backbones during sampling. Distribution matching distill…

  30. arXiv cs.CV TIER_1 English(EN) · Jiahao Zhan, Yan Wang, Yongrui Ma, Qunliang Xing, Ruchang Yao, Runtao Liu, Shijie Zhao, Tianfan Xue ·

    DuoMatching:用于少样本视频生成的联合边缘分布匹配

    arXiv:2610.03543v1 Announce Type: new Abstract: Streaming video generation has benefited from distribution matching distillation (DMD), which matches the joint distribution of video frames to a video teacher's approximation of the real video distribution. Although this joint matc…

  31. arXiv cs.CV TIER_1 English(EN) · Jiaxing Song, Weiqi Yan, You Huang, Mingte Qiu, Huazhong Liu, Xiaofeng Zhu, Yunshan Zhong ·

    TRAC:用于高效自回归视频生成的轨迹感知重用和自适应校正

    arXiv:2610.02779v1 Announce Type: new Abstract: In this paper, we present trajectory-aware reuse and adaptive correction (TRAC), a training-free framework for efficient autoregressive (AR) video generation. Existing acceleration methods mainly target single-trajectory generation …

  32. arXiv cs.CV TIER_1 English(EN) · Jeongwoo Shin, Youngyoon Choi, Sangwoo Jo, Hyunmog Kim, Sungjoon Choi, Joonseok Lee, Jaewoong Choi, Jaemoo Choi ·

    测试时用于长视频生成的分布内强制

    arXiv:2610.03120v1 Announce Type: new Abstract: Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works pr…

  33. arXiv cs.CV TIER_1 English(EN) · Yiwen Zhang, Haocheng Xi, Michael Tian-Yue Liu, Alexei A. Efros, Hadar Averbuch-Elor, Qianqian Wang, Haiwen Feng ·

    MosaiChunk:为自回归视频生成组合时空记忆

    arXiv:2610.02153v1 Announce Type: new Abstract: Long-horizon autoregressive video generation is limited by a finite context window. When an object or scene falls out of context, its fine-grained visual details may be lost and difficult to recover upon reappearance. To retain acce…

  34. arXiv cs.CV TIER_1 English(EN) · Mikhail Dereviannykh, Vikram Voleti, Simon Donne, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss ·

    SemanTok:用于高效自回归视频生成的可预测语义令牌

    arXiv:2610.00686v1 Announce Type: new Abstract: Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in ter…

  35. arXiv cs.CV TIER_1 English(EN) · Zhuo Ning, AmirHossein Naghi Razlighi, Sagi Polaczek, Daniel Cohen-Or, Ali Mahdavi-Amiri ·

    Soundwich:具有分层和可控音频的视频生成

    arXiv:2610.00691v1 Announce Type: new Abstract: Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workfl…

  36. arXiv cs.CV TIER_1 English(EN) · Jiho Jang, Jinyoung Kim, Nojun Kwak, Kyungjune Kim ·

    利用合成状态转换引导视频交互生成

    arXiv:2610.01039v1 Announce Type: new Abstract: While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting state transitions, a critical bottleneck for applications in robotics and VR/AR. To…

  37. arXiv cs.CV TIER_1 English(EN) · Tongcheng Zhang, Jun Zhu, Jianfei Chen ·

    面向视频生成中动态主题集上的主题一致性

    arXiv:2610.01052v1 Announce Type: new Abstract: We argue that as video generation extends to longer durations, subject consistency should be evaluated over \textit{dynamic subject sets}. We therefore introduce \textbf{DynSC-Eval}, an evaluation framework that dynamically tracks e…

  38. arXiv cs.CV TIER_1 English(EN) · Huanran Hu, Zihui Ren, Dingyi Yang, Zhinan Song, Guozheng Wu, Tiezheng Ge, Qin Jin ·

    DiVid:诊断视频生成模型中的维度特定多样性崩溃

    arXiv:2610.01661v1 Announce Type: new Abstract: Despite remarkable progress, video generation models often produce highly similar outputs when repeatedly sampled from the same prompt, limiting their usefulness for creative exploration. Existing diversity evaluations primarily rel…

  39. arXiv cs.CV TIER_1 English(EN) · Cusuh Ham, Fabian Caba Heilbron, Josef Sivic, Bryan Russell ·

    用户视频集中的记忆引导式B-roll生成

    arXiv:2610.01884v1 Announce Type: new Abstract: We introduce an approach for collection-grounded B-roll sequence generation. Given a user's video collection, a directive given in natural language, and a target duration, the goal is to produce a multi-shot sequence that complement…

  40. arXiv cs.CV TIER_1 English(EN) · Tahira Kazimi, Shubhankar Borse, Munawar Hayat, Fatih Porikli, Pinar Yanardag ·

    HiPhy:面向物理可信多原理视频生成的层级对齐

    arXiv:2610.02197v1 Announce Type: new Abstract: Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The pr…

  41. arXiv cs.CV TIER_1 English(EN) · Yuxin Cao, Wei Song, Shangzhi Xu, Jingling Xue, Jin Song Dong ·

    VideoSTF:对视频大语言模型输出重复性进行压力测试

    arXiv:2602.10639v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what models predict, leaving the stability of how they generate largely unexamined. We …

  42. arXiv cs.CV TIER_1 English(EN) · Zejing Rao, Ketong Ren, Xiaoqiang Liu, Yiping Meng, Guoxin Zhang, Fan Tang ·

    不惜成本:视频生成中流式提示切换的基于状态的过渡

    arXiv:2609.38691v1 Announce Type: new Abstract: Streaming video generators allow users to dynamically modulate video synthesis via mid-stream prompt switching. Existing streaming methods can respond to the updated instruction while still cutting corners, prematurely realizing goa…

  43. arXiv cs.CV TIER_1 English(EN) · Byoungwoo Park, Jaemoo Choi, Juho Lee, Yongxin Chen ·

    LongTake:学习在长时域视频生成中维持动态

    arXiv:2609.38562v1 Announce Type: new Abstract: World models, game simulators, and long-take video creation require coherent scene evolution and sustained dynamics over extended durations. Autoregressive (AR) video diffusion provides a natural framework for long-horizon generatio…

  44. arXiv cs.CV TIER_1 Italiano(IT) · Fangyu Lin, Xingtong Ge, Lunjie Zhu, Yi Zhang, Zhening Liu, Tianhang Wang, Mengfei Li, Yumeng Zhang, Guanglu Song, Yu Liu, Jun Zhang ·

    通过表示对抗蒸馏增强自回归视频生成

    arXiv:2609.40037v1 Announce Type: new Abstract: Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context and can propagate through subsequent rollouts, leading to detail degradation, struc…

  45. arXiv cs.CV TIER_1 English(EN) · Lingyu Liu, Yaxiong Wang, Li Zhu, Zhedong Zheng ·

    面向少样本视频生成的面向不确定性的自洽蒸馏

    arXiv:2609.39132v1 Announce Type: new Abstract: We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a co…

  46. arXiv cs.CV TIER_1 English(EN) · Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan ·

    FrameMorrow:面向长视频生成的未来引导帧选择与前瞻性Token

    arXiv:2609.38839v1 Announce Type: new Abstract: Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effec…

  47. arXiv cs.CV TIER_1 English(EN) · Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen ·

    LongLive-Plug:面向视频生成的一次性蒸馏

    arXiv:2609.38154v1 Announce Type: new Abstract: Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation…

  48. r/StableDiffusion TIER_2 English(EN) · /u/Total-Resort-3120 ·

    PDMD:面向视频扩散模型的预测分布匹配蒸馏

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1wxlyq2/pdmd_projected_distribution_matching_distillation/"> <img alt="PDMD: Projected Distribution Matching Distillation for Video Diffusion Models" src="https://external-preview.redd.it/OWZubGE1ZmFvaHRo…

  49. r/StableDiffusion TIER_2 English(EN) · /u/plsendfast ·

    面向视频生成模型的恒定大小内存!

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1ww9n5a/a_constantsize_memory_for_video_generative_models/"> <img alt="A constant-size memory for video generative models!" src="https://external-preview.redd.it/ZoNTItgscnVJIcJh8iOV7svC29KzUVsbZBRGcy79bN…