New benchmarks and models advance AI video generation quality and control · 10 sources tracked
ByPulseAugur Editorial·[28 sources]·
Recent research explores advancements in video generation, focusing on improving physical consistency, controllability, and efficiency. Papers introduce new benchmarks like FilmBench for cinematic quality and UniMoCa for unified motion and camera control. Models such as ContextMaster and VideoCoCo aim to handle complex, multi-shot video creation and physically consistent dynamics, respectively. Additionally, techniques like Token Radius Attention and MMPhysVideo are being developed to enhance the efficiency and physical plausibility of video generation models.
AI
IMPACT
These advancements aim to improve the realism, control, and efficiency of AI-generated videos, potentially impacting creative industries and content creation workflows.
RANK_REASON
The cluster contains multiple research papers introducing new models, benchmarks, and techniques for video generation.
arXiv:2512.11534v2 Announce Type: replace-cross Abstract: Key frame selection is essentially a set-level optimization problem: the quality of the selected subset depends on the interactions among frames, rather than the score of any single frame. Existing methods generally exhibi…
arXiv cs.AI
TIER_1English(EN)·Lingxiao Yang, Liu Liu, Moran Li, Han Feng, Wenjian Cao, Jiangning Zhang, Ye Shi·
arXiv:2608.05237v1 Announce Type: cross Abstract: Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean frames leak excessive local details, which causes …
Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typically rely on uniform sampling, or frame select…
arXiv:2504.01407v3 Announce Type: replace-cross Abstract: Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and the risk of losing essential context through naive downsampling. Inspired by the w…
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow …
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow …
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometr…
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introd…
arXiv:2607.24241v1 Announce Type: cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic …
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evalu…
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evalu…
arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level. We represent each frozen model--item pair by its res…
arXiv:2608.05707v1 Announce Type: new Abstract: Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reasoning demands…
arXiv:2608.05592v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typ…
arXiv:2608.05976v1 Announce Type: new Abstract: Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, and degrade when extrapolated to long videos, struggling to maintain long-range te…
arXiv:2608.04701v1 Announce Type: new Abstract: The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experience…
arXiv:2608.04956v1 Announce Type: new Abstract: Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, re…
arXiv:2608.02504v1 Announce Type: new Abstract: Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-level sparse methods share computation budgets across queries, overlooking token-sp…
arXiv cs.CV
TIER_1English(EN)·Shubo Lin, Xuanyang Zhang, Wei Cheng, Weiming Hu, Gang Yu, Jin Gao·
arXiv:2604.02817v2 Announce Type: replace Abstract: Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address this, we propose MMPhysVideo, the first study t…
arXiv:2608.01944v1 Announce Type: new Abstract: Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typi…
arXiv cs.CV
TIER_1English(EN)·Jian Xu, Yanning Wu, Delu Zeng, John Paisley, Qibin Zhao·
arXiv:2608.00617v1 Announce Type: new Abstract: Many physical attributes are \emph{irreversible}: ice melts but does not re-freeze, paper chars but does not un-burn. Do video generators respect this? We show the question is hard to measure, and that what can be measured reliably …
arXiv cs.CV
TIER_1English(EN)·Chong Gao, Jie Ma, Zhan Peng, Chongxiao Wang, Haoxue Wu, Jun Liang, Guanbin Li, Jing Li·
arXiv:2607.29545v1 Announce Type: new Abstract: Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, allowing diverse tasks to share complementary data and generative priors. Unifying…
arXiv:2607.27380v1 Announce Type: new Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text pro…
arXiv:2507.18046v2 Announce Type: replace Abstract: Recent advances in AI-generated video have shown strong performance on \emph{text-to-video} tasks, particularly for short clips depicting a single scene. However, current models struggle to generate longer videos with coherent s…
arXiv:2607.26529v1 Announce Type: new Abstract: Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended …