PulseAugur
实时 13:39:16
English(EN) Vidu S1: A Real-Time Interactive Video Generation Model

新研究通过改进的时间一致性和效率来增强视频生成

研究人员正在开发新的方法来改进视频生成模型,重点关注效率和时间一致性。一种方法,Hamiltonian Generative Networks (HGNs),旨在实现与帧率无关的连续时间预测,并提出了解决潜在幅度增长和截断误差累积等问题的修复方案。另一个重点领域是使用扩散模型统一视频生成和理解,Gen4U 等框架将生成表示重新用于视频分类和深度估计等任务。效率也通过诸如少步蒸馏和动态计算等技术得到解决,如 Dynamic-in-Few-StepOPSD-V 等模型所示,这些模型旨在在不牺牲质量的情况下降低计算成本。 AI

影响 视频生成模型的进步有望实现更高效、时间上更一致的内容创作,可能对媒体制作和模拟产生影响。

排序理由 arXiv 上发表了多篇研究论文,详细介绍了视频生成的新方法和模型。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 23 个来源。 我们如何撰写摘要 →

新研究通过改进的时间一致性和效率来增强视频生成

报道来源 [23]

  1. arXiv cs.AI TIER_1 English(EN) · Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu ·

    视频生成模型是通用的视觉学习者

    arXiv:2607.09024v1 Announce Type: cross Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper,…

  2. arXiv cs.LG TIER_1 English(EN) · Eli Laird, Corey Clark ·

    解锁汉密尔顿视频动力学模型中的时间泛化能力

    arXiv:2607.07763v1 Announce Type: new Abstract: World models are typically trained to predict discrete-time physical dynamics with a fixed step size baked into the model weights, preventing prediction at variable temporal resolutions. This matters for hierarchical planning, sim-t…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    视频生成模型是通用视觉学习者

    Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generat…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    OPSD-V:面向训练后少样本自回归视频生成器的策略内自蒸馏

    We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from error accumulation and weakened motion dynamics …

  5. arXiv cs.AI TIER_1 English(EN) · Yu Cheng, Siyue Yao, Zhongang Qi, Shanyan Guan, Wei Li, Fajie Yuan ·

    Dynamic-in-Few-Step:统一动态计算与少样本蒸馏以实现高效视频生成

    arXiv:2607.06631v1 Announce Type: cross Abstract: Video Diffusion Models (VDMs) have demonstrated superior generation quality but suffer from prohibitive computational costs. While recent few-step distillation techniques significantly accelerate inference, they typically enforce …

  6. arXiv cs.LG TIER_1 Deutsch(DE) · Michael King, Aravindh Mahendran, Matthew Koichi Grimes, Fedor Kitashov, Adham Elarabawy, Pedro Velez, Maks Ovsjanikov, Viorica P\u{a}tr\u{a}ucean ·

    Gen4U:通过扩散统一视频生成与理解

    arXiv:2607.06856v1 Announce Type: cross Abstract: Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematically probing …

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    OPSD-V:用于训练后少样本自回归视频生成器的策略内自蒸馏

    OPSD-V enhances few-step autoregressive video diffusion models by using real long-video data for temporal context during training, providing dense trajectory-level supervision that improves visual quality and motion dynamics without altering inference mechanisms.

  8. arXiv cs.AI TIER_1 English(EN) · Hengji Zhou, Lingxuan Huang, Jian Wang, Bing Zhou, Si Wu, Lianghao Xia, Chao Huang ·

    VideoAgent:视频理解与编辑一体化框架

    arXiv:2606.23327v2 Announce Type: replace-cross Abstract: Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two critical limitations: i) inability to handle …

  9. Hugging Face Daily Papers TIER_1 English(EN) ·

    CineMobile:用于电影级摄像机运动生成的设备端图像到视频扩散模型

    CineMobile enables efficient image-to-video generation on mobile devices through distillation-guided pruning, diffusion distillation, and hybrid quantization techniques while maintaining visual quality and achieving significant speedup.

  10. Hugging Face Daily Papers TIER_1 English(EN) ·

    Vidu S1:一个实时交互式视频生成模型

    Vidu S1 is a real-time interactive video generation model that supports voice-controlled digital character animation with infinite-length output and high frame rate on consumer hardware.

  11. arXiv cs.CV TIER_1 English(EN) · Thanh-Nhan Vo, Trong-Thuan Nguyen, Trung-Hoang Le, Tam V. Nguyen, Minh-Triet Tran ·

    SAGA:用于自回归视频生成的稳定加速引导

    arXiv:2607.08020v1 Announce Type: new Abstract: Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter, and struct…

  12. arXiv cs.CV TIER_1 English(EN) · Utkarsh A. Mishra, Yongxin Chen, Danfei Xu, Yang Liu, Xi Chen, Jiayuan Mao ·

    通过时间比率理解和缓解视频-动作泛化差距

    arXiv:2607.08127v1 Announce Type: new Abstract: Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on robotic action data. We refer to this discrepancy as t…

  13. arXiv cs.CV TIER_1 English(EN) · Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He, Yue Ma, Ziyu Wan, Yong Zhang, Xiaoming Wei, Qifeng Chen ·

    OPSD-V:用于训练后少样本自回归视频生成器的策略内自蒸馏

    arXiv:2607.08766v1 Announce Type: new Abstract: We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from …

  14. arXiv cs.CV TIER_1 English(EN) · Xavier Thomas, Youngsun Lim, Ananya Srinivasan, Audrey Zheng, Deepti Ghadiyaram ·

    生成式动作指示器:评估合成视频中的人类运动

    arXiv:2512.01803v3 Announce Type: replace Abstract: Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Mo…

  15. arXiv cs.CV TIER_1 English(EN) · Zixin Guo, Yehonathan Litman, Yifeng He, John Miller, Chuhan Chen, Deva Ramanan ·

    LightCrafter:用于可控一致性重照明的 PBR 条件视频扩散精炼

    arXiv:2607.08016v1 Announce Type: new Abstract: Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illu…

  16. arXiv cs.CV TIER_1 English(EN) · Cristian Sminchisescu ·

    视频生成模型是通用的视觉学习者

    Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generat…

  17. arXiv cs.CV TIER_1 English(EN) · Qifeng Chen ·

    OPSD-V:用于训练后少样本自回归视频生成器的策略内自蒸馏

    We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from error accumulation and weakened motion dynamics …

  18. arXiv cs.CV TIER_1 English(EN) · Jiayuan Mao ·

    通过时间比率理解和缓解视频-动作泛化差距

    Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on robotic action data. We refer to this discrepancy as the video-action generalization gap. In this pape…

  19. arXiv cs.CV TIER_1 English(EN) · Jintao Rong, Xin Xie, Xinyi Yu, Linlin Ou, Xinyu Zhang, Chunhua Shen, Dong Gong ·

    当蒸馏破坏运动控制:为快速视频生成器恢复生成轨迹

    arXiv:2506.19348v2 Announce Type: replace Abstract: Training-free motion customization imposes motion patterns from reference videos onto video generators through test-time computation. Most existing methods target full diffusion models, requiring many denoising steps and high co…

  20. arXiv cs.CV TIER_1 English(EN) · Minh-Triet Tran ·

    SAGA:用于自回归视频生成的稳定加速引导

    Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter, and structural drift. In this paper, we investigate this f…

  21. arXiv cs.CV TIER_1 English(EN) · Deva Ramanan ·

    LightCrafter:PBR 条件化视频扩散模型用于可控一致的重光照

    Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illumination. Existing methods follow two paradigms:…

  22. arXiv cs.CV TIER_1 Deutsch(DE) · Viorica Pătrăucean ·

    Gen4U:通过扩散统一视频生成与理解

    Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematically probing their intermediate activations using recent mutual…

  23. arXiv cs.CV TIER_1 English(EN) · Cong Wei, Quande Liu, Zixuan Ye, Qiulin Wang, Xintao Wang, Pengfei Wan, Kun Gai, Wenhu Chen ·

    UniVideo:视频的统一理解、生成和编辑

    arXiv:2510.08377v4 Announce Type: replace Abstract: Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified mo…