PulseAugur
实时 16:46:13

新基准和模型提升了 AI 视频生成质量和控制力 · 跟踪 10 个来源

近期研究探索了视频生成领域的进展,重点在于提高物理一致性、可控性和效率。论文介绍了新的基准,如用于电影级质量的 FilmBench 和用于统一运动与相机控制的 UniMoCaContextMasterVideoCoCo 等模型分别旨在处理复杂的多镜头视频创作和物理上一致的动态。此外,正在开发 Token Radius AttentionMMPhysVideo 等技术,以提高视频生成模型的效率和物理合理性。 AI

影响 这些进展旨在提高 AI 生成视频的真实感、控制力和效率,可能对创意产业和内容创作工作流程产生影响。

排序理由 该集群包含多篇研究论文,介绍了用于视频生成的新模型、基准和技术。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 28 个来源。 我们如何撰写摘要 →

新基准和模型提升了 AI 视频生成质量和控制力 · 跟踪 10 个来源

报道来源 [28]

  1. arXiv cs.CL TIER_1 English(EN) · Yiqing Yang, Yun Li, Daiqing Qi, Lehan Yang, Tianlong Wang, Wenhao Zhang, Sheng Li, Kin-man Lam ·

    HFS:用于高效视频理解的整体查询感知帧选择

    arXiv:2512.11534v2 Announce Type: replace-cross Abstract: Key frame selection is essentially a set-level optimization problem: the quality of the selected subset depends on the interactions among frames, rather than the score of any single frame. Existing methods generally exhibi…

  2. arXiv cs.AI TIER_1 English(EN) · Lingxiao Yang, Liu Liu, Moran Li, Han Feng, Wenjian Cao, Jiangning Zhang, Ye Shi ·

    上下文强制:揭示自回归视频扩散中的上下文效应

    arXiv:2608.05237v1 Announce Type: cross Abstract: Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean frames leak excessive local details, which causes …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    超越帧选择:用MLLMs重新思考长视频理解

    Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typically rely on uniform sampling, or frame select…

  4. arXiv cs.AI TIER_1 English(EN) · Yuan Zhang, Junwen Pan, Rui Zhang, Xin Wan, Qizhe Zhang, Ming Lu, Qi She, Shanghang Zhang ·

    ZoomV:用于高效长视频理解的时间缩放

    arXiv:2504.01407v3 Announce Type: replace-cross Abstract: Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and the risk of losing essential context through naive downsampling. Inspired by the w…

  5. arXiv cs.LG TIER_1 English(EN) · Zehua Chen, Junyou Wang, Yuxuan Jiang, Zhenying Fang, Yusheng Dai, Jianfei Chen, Ziwei Liu, Jun Zhu ·

    视觉表征至关重要:利用视频到音频生成中的时间差异

    arXiv:2608.04902v1 Announce Type: cross Abstract: Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis. However, existing conditional diffusion-based V2A methods typicall…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    ContextMaster:通过固定预算稀疏上下文路由实现交互式多镜头视频创作

    Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow …

  7. Hugging Face Daily Papers TIER_1 English(EN) ·

    ContextMaster:通过固定预算稀疏上下文路由实现交互式多镜头视频创作

    Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow …

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    UniWorld-View:通过视频扩散模型实现大基线视图合成

    The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometr…

  9. Hugging Face Daily Papers TIER_1 English(EN) ·

    VideoCoCo:通过代理双引擎系统实现物理一致性视频生成的代码即思维链

    Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introd…

  10. arXiv cs.AI TIER_1 English(EN) · Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi, Fei Ding, Weixu Qiao, Jinlin Wang, Xiaotong Lv, Peng Han, Zimeng Li, Fanshu Ding, Yushu Wang, Han Wu, Jingjing Chen, Chongxiao Wang, Yanhao Wu, Chenglong Huang, Xiaoqian Zhu, Jie Tian, Hua Li, Jingjing F… ·

    FilmBench:用于电影级视频生成的电影级基准测试

    arXiv:2607.24241v1 Announce Type: cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic …

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    FilmBench:用于电影级视频生成的电影级基准测试

    Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evalu…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    FilmBench:用于电影级视频生成的电影级基准测试

    Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evalu…

  13. arXiv cs.CV TIER_1 English(EN) · Wenzhang Sun, Chunfeng Wang, Xiangchen Yin, Yujia Chen, Hao Li, Kun Zhan ·

    Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs

    arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level. We represent each frozen model--item pair by its res…

  14. arXiv cs.CV TIER_1 English(EN) · Wang Chen, Yu Chen, Xiang Wang, Shuai Li, Jinfa Huang, Xiawu Zheng ·

    一个排名,任何预算:用于长视频理解的Matryoshka证据到上下文帧选择

    arXiv:2608.05707v1 Announce Type: new Abstract: Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reasoning demands…

  15. arXiv cs.CV TIER_1 English(EN) · Ziling Huang, Shin'ichi Satoh ·

    超越帧选择:用MLLMs重新思考长视频理解

    arXiv:2608.05592v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typ…

  16. arXiv cs.CV TIER_1 English(EN) · Haoning Yang, Xinyuan Chen, Yaohui Wang, Guo Lu ·

    Diff-VF:通过扩散模型实现无需训练的高质量长视频生成

    arXiv:2608.05976v1 Announce Type: new Abstract: Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, and degrade when extrapolated to long videos, struggling to maintain long-range te…

  17. arXiv cs.CV TIER_1 English(EN) · Haiyang Zhou, Wangbo Yu, Chaoran Feng, Xunyu Zhou, Yonghong Tian, Li Yuan ·

    UniWorld-View:通过视频扩散模型实现大基线视图合成

    arXiv:2608.04701v1 Announce Type: new Abstract: The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experience…

  18. arXiv cs.CV TIER_1 English(EN) · Xu Guo, Zhengxuan Wei, Xinghui Li, Hanzhuo Huang, Xinyu Liu, Xiangyang Luo, Min Wei, Yiran Zhu, Qiulin Wang, Yulong Xu, Xintao Wang, Pengfei Wan, Qi Fan, Xiangwang Hou ·

    ContextMaster:通过固定预算稀疏上下文路由实现交互式多镜头视频创作

    arXiv:2608.04956v1 Announce Type: new Abstract: Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, re…

  19. arXiv cs.CV TIER_1 English(EN) · Jiayu Chen, Zhikun Jiang, Maoliang Li, Jiayi Luo, Jiawei Yang, Zihao Zheng, Hengyi Zhang, Guojie Luo, Xiang Chen ·

    Token Radius Attention for Efficient Video Generation

    arXiv:2608.02504v1 Announce Type: new Abstract: Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-level sparse methods share computation budgets across queries, overlooking token-sp…

  20. arXiv cs.CV TIER_1 English(EN) · Shubo Lin, Xuanyang Zhang, Wei Cheng, Weiming Hu, Gang Yu, Jin Gao ·

    MMPhysVideo:通过联合RGB感知建模实现物理上可行的视频生成

    arXiv:2604.02817v2 Announce Type: replace Abstract: Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address this, we propose MMPhysVideo, the first study t…

  21. arXiv cs.CV TIER_1 English(EN) · Liming Tan, Ye Chen, Hao Zhang, Lirong Qian, Feifei Li, Bingbing Ni ·

    UniMoCa:将运动和相机控制统一为忠实人类视频生成的视觉代理

    arXiv:2608.01944v1 Announce Type: new Abstract: Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typi…

  22. arXiv cs.CV TIER_1 English(EN) · Jian Xu, Yanning Wu, Delu Zeng, John Paisley, Qibin Zhao ·

    视频生成中不可逆过程发育不良的诊断

    arXiv:2608.00617v1 Announce Type: new Abstract: Many physical attributes are \emph{irreversible}: ice melts but does not re-freeze, paper chars but does not un-burn. Do video generators respect this? We show the question is hard to measure, and that what can be measured reliably …

  23. arXiv cs.CV TIER_1 English(EN) · Chong Gao, Jie Ma, Zhan Peng, Chongxiao Wang, Haoxue Wu, Jun Liang, Guanbin Li, Jing Li ·

    MoRoute:动态路由用于上下文多模态视频生成

    arXiv:2607.29545v1 Announce Type: new Abstract: Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to share complementary data and generative priors. Unifying…

  24. arXiv cs.CV TIER_1 English(EN) · Haodong Li, Tianfei Ren, Xiaoxiao Ma, Chunmei Qing, Zhen Fang, Sipeng He, Ziyu Guo, Haoyu Wu, Juanxi Tian, Yihang Zou, Ruichuan An, Dongzhi Jiang, Boxue Yang, Ji Xie, Xu Huang, Wenhao Yan, Jialv Zou, Zhengrong Yue, Yaxin Luo, Xiaotong Li, Yuzhu Wang, Jun… ·

    VideoCoCo:通过代理双引擎系统实现物理一致视频生成的代码即推理(Code-as-CoT)

    arXiv:2607.27380v1 Announce Type: new Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text pro…

  25. arXiv cs.CV TIER_1 English(EN) · Hanwen Shen, Jiajie Lu, Yupeng Cao, Xiaonan Yang ·

    通过后训练增强视频生成中的场景过渡感知

    arXiv:2507.18046v2 Announce Type: replace Abstract: Recent advances in AI-generated video have shown strong performance on \emph{text-to-video} tasks, particularly for short clips depicting a single scene. However, current models struggle to generate longer videos with coherent s…

  26. arXiv cs.CV TIER_1 English(EN) · Yuyang Huang, Yabo Chen, Wenrui Dai, Ziyang Zheng, Haibin Huang, Chi Zhang, Junni Zou, Hongkai Xiong, Xuelong Li ·

    CineWeaver:无需训练即可进行参考控制的多镜头长视频生成,用于电影叙事

    arXiv:2607.26529v1 Announce Type: new Abstract: Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended …

  27. r/StableDiffusion TIER_2 English(EN) · /u/Total-Resort-3120 ·

    Sol-Attn:通过即时注意力稀疏化加速视频生成推理。

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vdsuz4/solattn_accelerating_video_generation_inference/"> <img alt="Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification." src="https://external-preview.redd.it/czVsM…

  28. r/StableDiffusion TIER_2 English(EN) · /u/Sporeboss ·

    MiniMax H3:下一代开放权重多模态视频生成模型

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1vb96vo/minimax_h3_the_nextgen_openweight_multimodal/"> <img alt="MiniMax H3: The Next-Gen Open-Weight Multimodal Video Generation Model" src="https://external-preview.redd.it/Va-n_NvWCn6A2wdVoZEwe3hQX-BF…