PulseAugur
EN
LIVE 11:46:39

New research enhances video generation with improved temporal consistency and efficiency

Researchers are developing new methods to improve video generation models, focusing on efficiency and temporal consistency. One approach, Hamiltonian Generative Networks (HGNs), aims for continuous-time prediction independent of frame rates, with proposed fixes for issues like latent magnitude growth and truncation error accumulation. Another area of focus is unifying video generation and understanding using diffusion models, with frameworks like Gen4U repurposing generative representations for tasks such as video classification and depth estimation. Efficiency is also being addressed through techniques like few-step distillation and dynamic computation, as seen in models like Dynamic-in-Few-Step and OPSD-V, which aim to reduce computational costs without sacrificing quality. AI

IMPACT Advances in video generation models promise more efficient and temporally consistent content creation, potentially impacting media production and simulation.

RANK_REASON Multiple research papers published on arXiv detailing new methods and models for video generation.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 23 sources. How we write summaries →

New research enhances video generation with improved temporal consistency and efficiency

COVERAGE [23]

  1. arXiv cs.AI TIER_1 English(EN) · Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu ·

    Video Generation Models are General-Purpose Vision Learners

    arXiv:2607.09024v1 Announce Type: cross Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper,…

  2. arXiv cs.LG TIER_1 English(EN) · Eli Laird, Corey Clark ·

    Unlocking Temporal Generalization in Hamiltonian Video Dynamics Models

    arXiv:2607.07763v1 Announce Type: new Abstract: World models are typically trained to predict discrete-time physical dynamics with a fixed step size baked into the model weights, preventing prediction at variable temporal resolutions. This matters for hierarchical planning, sim-t…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    Video Generation Models are General-Purpose Vision Learners

    Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generat…

  4. Hugging Face Daily Papers TIER_1 English(EN) ·

    OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

    We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from error accumulation and weakened motion dynamics …

  5. arXiv cs.AI TIER_1 English(EN) · Yu Cheng, Siyue Yao, Zhongang Qi, Shanyan Guan, Wei Li, Fajie Yuan ·

    Dynamic-in-Few-Step: Unifying Dynamic Computation and Few-Step Distillation for Efficient Video Generation

    arXiv:2607.06631v1 Announce Type: cross Abstract: Video Diffusion Models (VDMs) have demonstrated superior generation quality but suffer from prohibitive computational costs. While recent few-step distillation techniques significantly accelerate inference, they typically enforce …

  6. arXiv cs.LG TIER_1 Deutsch(DE) · Michael King, Aravindh Mahendran, Matthew Koichi Grimes, Fedor Kitashov, Adham Elarabawy, Pedro Velez, Maks Ovsjanikov, Viorica P\u{a}tr\u{a}ucean ·

    Gen4U: Unifying Video Generation and Understanding via Diffusion

    arXiv:2607.06856v1 Announce Type: cross Abstract: Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematically probing …

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

    OPSD-V enhances few-step autoregressive video diffusion models by using real long-video data for temporal context during training, providing dense trajectory-level supervision that improves visual quality and motion dynamics without altering inference mechanisms.

  8. arXiv cs.AI TIER_1 English(EN) · Hengji Zhou, Lingxuan Huang, Jian Wang, Bing Zhou, Si Wu, Lianghao Xia, Chao Huang ·

    VideoAgent: All-in-One Framework for Video Understanding and Editing

    arXiv:2606.23327v2 Announce Type: replace-cross Abstract: Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two critical limitations: i) inability to handle …

  9. Hugging Face Daily Papers TIER_1 English(EN) ·

    CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

    CineMobile enables efficient image-to-video generation on mobile devices through distillation-guided pruning, diffusion distillation, and hybrid quantization techniques while maintaining visual quality and achieving significant speedup.

  10. Hugging Face Daily Papers TIER_1 English(EN) ·

    Vidu S1: A Real-Time Interactive Video Generation Model

    Vidu S1 is a real-time interactive video generation model that supports voice-controlled digital character animation with infinite-length output and high frame rate on consumer hardware.

  11. arXiv cs.CV TIER_1 English(EN) · Thanh-Nhan Vo, Trong-Thuan Nguyen, Trung-Hoang Le, Tam V. Nguyen, Minh-Triet Tran ·

    SAGA: Stable Acceleration Guidance for Autoregressive Video Generation

    arXiv:2607.08020v1 Announce Type: new Abstract: Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter, and struct…

  12. arXiv cs.CV TIER_1 English(EN) · Utkarsh A. Mishra, Yongxin Chen, Danfei Xu, Yang Liu, Xi Chen, Jiayuan Mao ·

    Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

    arXiv:2607.08127v1 Announce Type: new Abstract: Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on robotic action data. We refer to this discrepancy as t…

  13. arXiv cs.CV TIER_1 English(EN) · Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He, Yue Ma, Ziyu Wan, Yong Zhang, Xiaoming Wei, Qifeng Chen ·

    OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

    arXiv:2607.08766v1 Announce Type: new Abstract: We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from …

  14. arXiv cs.CV TIER_1 English(EN) · Xavier Thomas, Youngsun Lim, Ananya Srinivasan, Audrey Zheng, Deepti Ghadiyaram ·

    Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

    arXiv:2512.01803v3 Announce Type: replace Abstract: Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Mo…

  15. arXiv cs.CV TIER_1 English(EN) · Zixin Guo, Yehonathan Litman, Yifeng He, John Miller, Chuhan Chen, Deva Ramanan ·

    LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting

    arXiv:2607.08016v1 Announce Type: new Abstract: Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illu…

  16. arXiv cs.CV TIER_1 English(EN) · Cristian Sminchisescu ·

    Video Generation Models are General-Purpose Vision Learners

    Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generat…

  17. arXiv cs.CV TIER_1 English(EN) · Qifeng Chen ·

    OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

    We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from error accumulation and weakened motion dynamics …

  18. arXiv cs.CV TIER_1 English(EN) · Jiayuan Mao ·

    Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

    Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on robotic action data. We refer to this discrepancy as the video-action generalization gap. In this pape…

  19. arXiv cs.CV TIER_1 English(EN) · Jintao Rong, Xin Xie, Xinyi Yu, Linlin Ou, Xinyu Zhang, Chunhua Shen, Dong Gong ·

    When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators

    arXiv:2506.19348v2 Announce Type: replace Abstract: Training-free motion customization imposes motion patterns from reference videos onto video generators through test-time computation. Most existing methods target full diffusion models, requiring many denoising steps and high co…

  20. arXiv cs.CV TIER_1 English(EN) · Minh-Triet Tran ·

    SAGA: Stable Acceleration Guidance for Autoregressive Video Generation

    Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter, and structural drift. In this paper, we investigate this f…

  21. arXiv cs.CV TIER_1 English(EN) · Deva Ramanan ·

    LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting

    Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illumination. Existing methods follow two paradigms:…

  22. arXiv cs.CV TIER_1 Deutsch(DE) · Viorica Pătrăucean ·

    Gen4U: Unifying Video Generation and Understanding via Diffusion

    Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematically probing their intermediate activations using recent mutual…

  23. arXiv cs.CV TIER_1 English(EN) · Cong Wei, Quande Liu, Zixuan Ye, Qiulin Wang, Xintao Wang, Pengfei Wan, Kun Gai, Wenhu Chen ·

    UniVideo: Unified Understanding, Generation, and Editing for Videos

    arXiv:2510.08377v4 Announce Type: replace Abstract: Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified mo…