PulseAugur
EN
LIVE 14:13:46

New benchmarks and models advance AI video generation quality and control · 10 sources tracked

Recent research explores advancements in video generation, focusing on improving physical consistency, controllability, and efficiency. Papers introduce new benchmarks like FilmBench for cinematic quality and UniMoCa for unified motion and camera control. Models such as ContextMaster and VideoCoCo aim to handle complex, multi-shot video creation and physically consistent dynamics, respectively. Additionally, techniques like Token Radius Attention and MMPhysVideo are being developed to enhance the efficiency and physical plausibility of video generation models. AI

IMPACT These advancements aim to improve the realism, control, and efficiency of AI-generated videos, potentially impacting creative industries and content creation workflows.

RANK_REASON The cluster contains multiple research papers introducing new models, benchmarks, and techniques for video generation.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 28 sources. How we write summaries →

New benchmarks and models advance AI video generation quality and control · 10 sources tracked

COVERAGE [28]

  1. arXiv cs.CL TIER_1 English(EN) · Yiqing Yang, Yun Li, Daiqing Qi, Lehan Yang, Tianlong Wang, Wenhao Zhang, Sheng Li, Kin-man Lam ·

    HFS: Holistic Query-Aware Frame Selection for Efficient Video Understanding

    arXiv:2512.11534v2 Announce Type: replace-cross Abstract: Key frame selection is essentially a set-level optimization problem: the quality of the selected subset depends on the interactions among frames, rather than the score of any single frame. Existing methods generally exhibi…

  2. arXiv cs.AI TIER_1 English(EN) · Lingxiao Yang, Liu Liu, Moran Li, Han Feng, Wenjian Cao, Jiangning Zhang, Ye Shi ·

    In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion

    arXiv:2608.05237v1 Announce Type: cross Abstract: Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean frames leak excessive local details, which causes …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

    Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typically rely on uniform sampling, or frame select…

  4. arXiv cs.AI TIER_1 English(EN) · Yuan Zhang, Junwen Pan, Rui Zhang, Xin Wan, Qizhe Zhang, Ming Lu, Qi She, Shanghang Zhang ·

    ZoomV: Temporal Zoom-in for Efficient Long Video Understanding

    arXiv:2504.01407v3 Announce Type: replace-cross Abstract: Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and the risk of losing essential context through naive downsampling. Inspired by the w…

  5. arXiv cs.LG TIER_1 English(EN) · Zehua Chen, Junyou Wang, Yuxuan Jiang, Zhenying Fang, Yusheng Dai, Jianfei Chen, Ziwei Liu, Jun Zhu ·

    Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

    arXiv:2608.04902v1 Announce Type: cross Abstract: Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis. However, existing conditional diffusion-based V2A methods typicall…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

    Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow …

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

    Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow …

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

    The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometr…

  9. Hugging Face Daily Papers TIER_1 English(EN) ·

    VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

    Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introd…

  10. arXiv cs.AI TIER_1 English(EN) · Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing F… ·

    FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

    arXiv:2607.24241v1 Announce Type: cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic …

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

    Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evalu…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

    Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evalu…

  13. arXiv cs.CV TIER_1 English(EN) · Wenzhang Sun, Chunfeng Wang, Xiangchen Yin, Yujia Chen, Hao Li, Kun Zhan ·

    Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs

    arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level. We represent each frozen model--item pair by its res…

  14. arXiv cs.CV TIER_1 English(EN) · Wang Chen, Yu Chen, Xiang Wang, Shuai Li, Jinfa Huang, Xiawu Zheng ·

    One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

    arXiv:2608.05707v1 Announce Type: new Abstract: Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reasoning demands…

  15. arXiv cs.CV TIER_1 English(EN) · Ziling Huang, Shin'ichi Satoh ·

    Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

    arXiv:2608.05592v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typ…

  16. arXiv cs.CV TIER_1 English(EN) · Haoning Yang, Xinyuan Chen, Yaohui Wang, Guo Lu ·

    Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model

    arXiv:2608.05976v1 Announce Type: new Abstract: Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, and degrade when extrapolated to long videos, struggling to maintain long-range te…

  17. arXiv cs.CV TIER_1 English(EN) · Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou, Yonghong Tian, Li Yuan ·

    UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

    arXiv:2608.04701v1 Announce Type: new Abstract: The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experience…

  18. arXiv cs.CV TIER_1 English(EN) · Xu Guo, Zhengxuan Wei, Xinghui Li, Hanzhuo Huang, Xinyu Liu, Xiangyang Luo, Min Wei, Yiran Zhu, Qiulin Wang, Yulong Xu, Xintao Wang, Pengfei Wan, Qi Fan, Xiangwang Hou ·

    ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

    arXiv:2608.04956v1 Announce Type: new Abstract: Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, re…

  19. arXiv cs.CV TIER_1 English(EN) · Jiayu Chen, Zhikun Jiang, Maoliang Li, Jiayi Luo, Jiawei Yang, Zihao Zheng, Hengyi Zhang, Guojie Luo, Xiang Chen ·

    Token Radius Attention for Efficient Video Generation

    arXiv:2608.02504v1 Announce Type: new Abstract: Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-level sparse methods share computation budgets across queries, overlooking token-sp…

  20. arXiv cs.CV TIER_1 English(EN) · Shubo Lin, Xuanyang Zhang, Wei Cheng, Weiming Hu, Gang Yu, Jin Gao ·

    MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling

    arXiv:2604.02817v2 Announce Type: replace Abstract: Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address this, we propose MMPhysVideo, the first study t…

  21. arXiv cs.CV TIER_1 English(EN) · Liming Tan, Ye Chen, Hao Zhang, Lirong Qian, Feifei Li, Bingbing Ni ·

    UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

    arXiv:2608.01944v1 Announce Type: new Abstract: Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typi…

  22. arXiv cs.CV TIER_1 English(EN) · Jian Xu, Yanning Wu, Delu Zeng, John Paisley, Qibin Zhao ·

    Diagnosing Under-Development of Irreversible Processes in Video Generation

    arXiv:2608.00617v1 Announce Type: new Abstract: Many physical attributes are \emph{irreversible}: ice melts but does not re-freeze, paper chars but does not un-burn. Do video generators respect this? We show the question is hard to measure, and that what can be measured reliably …

  23. arXiv cs.CV TIER_1 English(EN) · Chong Gao, Jie Ma, Zhan Peng, Chongxiao Wang, Haoxue Wu, Jun Liang, Guanbin Li, Jing Li ·

    MoRoute: Dynamic Routing for In-Context Multimodal Video Generation

    arXiv:2607.29545v1 Announce Type: new Abstract: Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to share complementary data and generative priors. Unifying…

  24. arXiv cs.CV TIER_1 English(EN) · Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Jun… ·

    VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

    arXiv:2607.27380v1 Announce Type: new Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text pro…

  25. arXiv cs.CV TIER_1 English(EN) · Hanwen Shen, Jiajie Lu, Yupeng Cao, Xiaonan Yang ·

    Enhancing Scene Transition Awareness in Video Generation via Post-Training

    arXiv:2507.18046v2 Announce Type: replace Abstract: Recent advances in AI-generated video have shown strong performance on \emph{text-to-video} tasks, particularly for short clips depicting a single scene. However, current models struggle to generate longer videos with coherent s…

  26. arXiv cs.CV TIER_1 English(EN) · Yuyang Huang, Yabo Chen, Wenrui Dai, Ziyang Zheng, Haibin Huang, Chi Zhang, Junni Zou, Hongkai Xiong, Xuelong Li ·

    CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

    arXiv:2607.26529v1 Announce Type: new Abstract: Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended …

  27. r/StableDiffusion TIER_2 English(EN) · /u/Total-Resort-3120 ·

    Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification.

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vdsuz4/solattn_accelerating_video_generation_inference/"> <img alt="Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification." src="https://external-preview.redd.it/czVsM…

  28. r/StableDiffusion TIER_2 English(EN) · /u/Sporeboss ·

    MiniMax H3: The Next-Gen Open-Weight Multimodal Video Generation Model

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vb96vo/minimax_h3_the_nextgen_openweight_multimodal/"> <img alt="MiniMax H3: The Next-Gen Open-Weight Multimodal Video Generation Model" src="https://external-preview.redd.it/Va-n_NvWCn6A2wdVoZEwe3hQX-BF…