PulseAugur
EN
LIVE 14:59:02

New research tackles video generation challenges in control, consistency, and efficiency

Researchers are developing advanced techniques for video generation, focusing on improving control, consistency, and efficiency. CineOrchestra aims to unify control over subjects, cameras, and shot transitions in cinematic videos. TetherCache addresses drift and quality degradation in long-form autoregressive video generation by managing cache memory. Argus enhances subject preservation across various challenging conditions using a novel identity injection method. MilliVid employs a hierarchical latent space for long-range consistency, while RhymeFlow accelerates diffusion transformers by decoupling denoising trajectories. Echo-Infinity introduces learnable evolving memory for real-time infinite video generation, and MBench provides a benchmark for evaluating memory capabilities in video world models. AI

IMPACT Advancements in video generation models are improving control, consistency, and efficiency, paving the way for more sophisticated applications.

RANK_REASON Multiple research papers published on arXiv and Hugging Face detailing new methods and benchmarks for video generation.

Read on Hugging Face Daily Papers →

AI-generated summary · Google Gemini · from 59 sources. How we write summaries →

New research tackles video generation challenges in control, consistency, and efficiency

COVERAGE [59]

  1. arXiv cs.AI TIER_1 English(EN) · Sixiao Zheng, Zimian Peng, Yanpeng Zhou, Yi Zhu, Hang Xu, Xiangru Huang, Yanwei Fu ·

    VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation

    arXiv:2502.07531v5 Announce Type: replace-cross Abstract: Controllable image-to-video (I2V) generation transforms a reference image into a coherent video guided by user-specified control signals. While precise control over camera motion, object motion, and lighting is essential f…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation

    Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal generation, whereas practical video creation often requires flexi…

  3. arXiv cs.AI TIER_1 English(EN) · Zhenyu Yang, Kairui Zhang, Bing Wang, Shengsheng Qian, Changsheng Xu ·

    LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams

    arXiv:2606.17798v1 Announce Type: cross Abstract: Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preserve long-ho…

  4. arXiv cs.LG TIER_1 English(EN) · Adnan El Assadi, Roman Solomatin, Isaac Chung, Chenghao Xiao, Deep Shah, Manan Dey, Shriya Sudhakar, Zacharie Bugaud, Wissam Siblini, Ayush Sunil Munot, Yashwanth Devavarapu, Rakshitha Ireddi, Michelle Yang, M\'arton Kardos, Niklas Muennighoff, Kenneth E… ·

    MVEB: Massive Video Embedding Benchmark

    arXiv:2606.14958v1 Announce Type: cross Abstract: We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classification, retrieval, and video-centric question answerin…

  5. arXiv cs.AI TIER_1 English(EN) · Sharath Girish, Tsai-Shien Chen, Zhikang Dong, Mukesh Singhal, Hao Chen, Sergey Tulyakov, Aliaksandr Siarohin ·

    CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation

    arXiv:2606.13768v1 Announce Type: cross Abstract: Cinematic video depicts multiple subjects acting or interacting at specific moments, captured with deliberate camera movement, and stitched together by shot transitions. Together, these elements demand a level of fine-grained cont…

  6. Hugging Face Daily Papers TIER_1 English(EN) ·

    PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

    PermaVid addresses long-term video consistency after edits by using multi-modal memory banks that separate appearance and geometric structure, enabling coherent video generation across time and viewpoints.

  7. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Kenneth Enevoldsen ·

    MVEB: Massive Video Embedding Benchmark

    We introduce the Massive Video Embedding Benchmark (MVEB), a 23-task benchmark for video embeddings spanning classification, zero-shot classification, clustering, pair classification, retrieval, and video-centric question answering. We evaluate 33 models and find that no single m…

  8. arXiv cs.AI TIER_1 English(EN) · Yu Meng, Xiangyang Luo, Letian Li, Wenyuan Jiang, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang ·

    TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment

    arXiv:2606.13035v1 Announce Type: cross Abstract: Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minu…

  9. Hugging Face Daily Papers TIER_1 English(EN) ·

    MVEB: Massive Video Embedding Benchmark

    A large-scale video embedding benchmark evaluates diverse models across multiple video understanding tasks, revealing that different model architectures excel in specific domains and demonstrating the nuanced impact of audio on performance based on dataset characteristics.

  10. Hugging Face Daily Papers TIER_1 English(EN) ·

    Memento: Reconstruct to Remember for Consistent Long Video Generation

    Memento is a subject-reconstruction-guided framework that improves long-form video generation by preserving recurring subjects through memory-based reconstruction and dual-query mechanisms.

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment

    Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limit…

  12. arXiv cs.AI TIER_1 English(EN) · Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan ·

    ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

    arXiv:2606.11670v1 Announce Type: cross Abstract: Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts a…

  13. arXiv cs.LG TIER_1 English(EN) · Ishaan Preetam Chandratreya, David Charatan, Basile Van Hoorick, Sergey Zakharov, Vitor Guizilini, Phillip Isola, Vincent Sitzmann ·

    MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation

    arXiv:2606.09056v1 Announce Type: cross Abstract: Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames require impractically long transformer sequence lengths. We show that this issue …

  14. Hugging Face Daily Papers TIER_1 English(EN) ·

    MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

    A new benchmark called MBench is introduced to evaluate the memory capabilities of video world models, focusing on entity, environment, and causal consistency over extended temporal horizons.

  15. Hugging Face Daily Papers TIER_1 English(EN) ·

    MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation

    Video generative models achieve improved long-range consistency through coarse-to-fine token generation using a multi-scale autoencoder and diffusion model architecture.

  16. Hugging Face Daily Papers TIER_1 English(EN) ·

    LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

    LoomVideo presents an efficient 5B-parameter unified architecture for video generation and editing that reduces computational overhead through novel conditioning mechanisms and multi-modal alignment techniques.

  17. Hugging Face Daily Papers TIER_1 English(EN) ·

    RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling

    RhymeFlow accelerates diffusion transformers for video generation by decoupling denoising trajectories across frames, using keyframe anchoring and latent trajectory projection to maintain visual quality while reducing computational overhead.

  18. Hugging Face Daily Papers TIER_1 English(EN) ·

    Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation

    Echo Infinity enables real-time infinite video generation using learnable evolving memory and unified relative RoPE to overcome limitations in existing autoregressive methods.

  19. arXiv cs.AI TIER_1 English(EN) · Chenxu Wang, Mingda Chen ·

    Knowledge-Intensive Video Generation

    arXiv:2606.01285v1 Announce Type: cross Abstract: Text-to-video generation has advanced rapidly in visual quality, but remains under-evaluated for factuality and practical usefulness. We introduce knowledge-intensive video generation (KIVI), where models generate videos from shor…

  20. Hugging Face Daily Papers TIER_1 English(EN) ·

    LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

    LongLive-RAG addresses long-video generation challenges by using retrieval-augmented generation to overcome error accumulation from sliding-window attention, enabling better temporal coherence and quality.

  21. arXiv cs.CV TIER_1 English(EN) · Siyi Chen, Shaowei Liu, Yixuan Jia, Zian Wang, Huan Ling, Qing Qu, Jun Gao ·

    Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation

    arXiv:2606.18478v2 Announce Type: replace Abstract: Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matching Distillation (DMD) and its successor DMD2 achieved strong generation quality a…

  22. arXiv cs.CV TIER_1 English(EN) · Hui Ren, Yuval Alaluf, Omer Bar Tal, Alexander Schwing, Antonio Torralba, Yael Vinker ·

    VideoSketcher: Sequential Sketch Generation Using Video Model Priors

    arXiv:2602.15819v2 Announce Type: replace Abstract: Sketching is inherently sequential: strokes are drawn progressively to explore and refine ideas. Yet most generative approaches treat sketches as static images, ignoring the temporal process underlying creative exploration. Mode…

  23. arXiv cs.CV TIER_1 English(EN) · Yang Tan, Junlong Tong, Linan Yue, Hao Wu, Pengfei Fang, Xiaoyu Shen ·

    ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference

    arXiv:2606.19849v1 Announce Type: new Abstract: Streaming VideoLLMs must continuously process incoming video while maintaining low query latency, making both video-ingestion throughput and query-time responsiveness critical for real-time deployment. Existing methods largely focus…

  24. arXiv cs.CV TIER_1 English(EN) · Yin Li ·

    UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation

    Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal generation, whereas practical video creation often requires flexi…

  25. arXiv cs.CV TIER_1 English(EN) · Zhe Zhao ·

    Bridging Creative Intent and Visual Quality: Creator-Driven Recurrent Video Generation with Agentic Feedback Loops

    Generative AI has made content creation increasingly accessible, but many AI-generated videos lack narrative coherence and creative direction, issues that become more substantial at longer durations. Unlike coding, where AI generation benefits from reliable feedback and technique…

  26. arXiv cs.CV TIER_1 English(EN) · Jun Gao ·

    Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation

    Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matching Distillation (DMD) and its successor DMD2 achieved strong generation quality and fast convergence. However, due to the nature of t…

  27. arXiv cs.CV TIER_1 English(EN) · Changsheng Xu ·

    LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams

    Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preserve long-horizon contextual memory. These obstacles undermine…

  28. arXiv cs.CV TIER_1 English(EN) · Yizhou Zhao, Yifan Wang, Xiaoyuan Wang, Yushu Wu, Hao Zhang, Moayed Haji-Ali, Rameen Abdal, Ashkan Mirzaei, Yanyu Li, Willi Menapace, Laszlo Jeni, Sergey Tulyakov, Peter Wonka, Chaoyang Wang ·

    GeoStream: Toward Precise Camera Controlled Streaming Video Generation

    arXiv:2606.15162v1 Announce Type: new Abstract: Accurate interactive camera control is essential for video-based world models, but most existing approaches learn camera motion implicitly, leading to inaccurate control under out-of-distribution trajectories. Explicit geometric con…

  29. arXiv cs.CV TIER_1 English(EN) · Xinlei Yin, Xiulian Peng, Xiao Li, Zhiwei Xiong, Yan Lu ·

    Closed-Loop Triplet Synergistic Generation for Long-Form Video

    arXiv:2606.16184v1 Announce Type: new Abstract: Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. While storyboard-driven pipelines improve controllability, they are often executed in a feed-forward manne…

  30. arXiv cs.CV TIER_1 English(EN) · Shuai Yang, Bingjie Gao, Ziwei Liu, Jiaqi Wang, Dahua Lin, Tong Wu ·

    PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

    arXiv:2606.16449v1 Announce Type: new Abstract: Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time and viewpoints. However, existing memory designs stru…

  31. arXiv cs.CV TIER_1 English(EN) · Chaoyu Li, Tianzhi Li, Fei Tao, Zhenyu Zhao, Ziqian Wu, Maozheng Zhao, Juntong Song, Cheng Niu, Pooyan Fazli ·

    FrameOracle: Learning What to See and How Much to See in Videos

    arXiv:2510.03584v3 Announce Type: replace Abstract: Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such …

  32. arXiv cs.CV TIER_1 English(EN) · Tong Wu ·

    PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

    Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time and viewpoints. However, existing memory designs struggle to maintain long-term consistency after suc…

  33. arXiv cs.CV TIER_1 English(EN) · Xuan Wei, Longbin Ji, Guan Wang, Xiangrui Liu, Zhenyu Zhang, Shuohuan Wang, Yu Sun, Qingqi Hong ·

    Memento: Reconstruct to Remember for Consistent Long Video Generation

    arXiv:2606.14667v1 Announce Type: new Abstract: Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating videos shot by s…

  34. arXiv cs.CV TIER_1 English(EN) · Qingqi Hong ·

    Memento: Reconstruct to Remember for Consistent Long Video Generation

    Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating videos shot by shot. However, they mainly focus on optimizing pl…

  35. arXiv cs.CV TIER_1 English(EN) · Xiao-Ping Zhang ·

    TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment

    Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limit…

  36. arXiv cs.CV TIER_1 English(EN) · Jingxu Zhang, Yuqian Hong, Daneul Kim, Kai Qiu, Qi Dai, Jianmin Bao, Yifan Yang, Xiaoyan Sun, Chong Luo ·

    A Comprehensive Ecosystem for Open-Domain Customized Video Generation

    arXiv:2606.11783v1 Announce Type: new Abstract: Recent progress in video generation has shown impressive visual synthesis capabilities. However, open-domain customized video generation remains limited by the lack of large-scale, annotated datasets capturing diverse identity-speci…

  37. arXiv cs.CV TIER_1 English(EN) · Chong Luo ·

    A Comprehensive Ecosystem for Open-Domain Customized Video Generation

    Recent progress in video generation has shown impressive visual synthesis capabilities. However, open-domain customized video generation remains limited by the lack of large-scale, annotated datasets capturing diverse identity-specific attributes. To address this, we introduce Pe…

  38. arXiv cs.CV TIER_1 English(EN) · Cong Wang, Zhentao Yu, Hongmei Wang, Weicong Liang, Zixiang Zhou, Zilin Yang, Jiarong Ou, Rui Chen, Yuan Zhou, Qinglin Lu ·

    HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation

    arXiv:2606.10839v1 Announce Type: new Abstract: Current identity-consistent video generation methods struggle to preserve appearance fidelity under large viewpoint changes. While introducing multi-view reference input offers a natural solution, progress remains constrained by the…

  39. arXiv cs.CV TIER_1 English(EN) · Qinglin Lu ·

    HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation

    Current identity-consistent video generation methods struggle to preserve appearance fidelity under large viewpoint changes. While introducing multi-view reference input offers a natural solution, progress remains constrained by the lack of effective frameworks for multi-view inp…

  40. arXiv cs.CV TIER_1 English(EN) · Xinshuang Liu, Runfa Blark Li, Truong Nguyen ·

    Consistency-Preserving Diverse Video Generation

    arXiv:2602.15287v2 Announce Type: replace Abstract: Text-to-video generation is expensive, so only a few samples are typically produced per prompt. In this low-sample regime, maximizing the value of each batch requires high cross-video diversity. Recent methods improve diversity …

  41. arXiv cs.CV TIER_1 English(EN) · Tao Liu, Leela Krishna, Gouti Pavan Kumar, Sreeja K, Vishav Garg ·

    V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation

    arXiv:2606.05665v1 Announce Type: new Abstract: Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence with the source video, which existing T2V and I2V metrics do not capture. We intr…

  42. arXiv cs.CV TIER_1 English(EN) · Jianzong Wu, Hao Lian, Jiongfan Yang, Dachao Hao, Ye Tian, Yunhai Tong, Jingyuan Zhu, Biaolong Chen, Qiaosong Qi, Aixi Zhang, Wanggui He, Mushui Liu, Jinlong Liu, Hao Jiang ·

    LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

    arXiv:2606.06042v1 Announce Type: new Abstract: Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier field. Existing unified frameworks predominantly rely on massive models (typically …

  43. arXiv cs.CV TIER_1 English(EN) · Hao Jiang ·

    LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

    Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier field. Existing unified frameworks predominantly rely on massive models (typically 13B parameters or more) and incorporate source v…

  44. arXiv cs.CV TIER_1 English(EN) · Yuxuan Bian, Zeyue Xue, Songchun Zhang, Shiyi Zhang, Weiyang Jin, Yaowei Li, Junhao Zhuang, Haoran Li, Jie Huang, Haoyang Huang, Nan Duan, Qiang Xu ·

    Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation

    arXiv:2606.04527v1 Announce Type: cross Abstract: We present Echo Infinity, an autoregressive (AR) framework towards real-time infinite video generation that employs a learnable evolving memory to dynamically filter, abstract, and compress any-length history at constant cost. Exi…

  45. arXiv cs.CV TIER_1 English(EN) · Qiang Xu ·

    Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation

    We present Echo Infinity, an autoregressive (AR) framework towards real-time infinite video generation that employs a learnable evolving memory to dynamically filter, abstract, and compress any-length history at constant cost. Existing methods mainly curate memory with predefined…

  46. arXiv cs.CV TIER_1 English(EN) · Zhengxuan Wei, Xu Guo, Xinghui Li, Xunzhi Xiang, Min Wei, Yiran Zhu, Qiulin Wang, Xintao Wang, Pengfei Wan, Xiangwang Hou, Qi Fan ·

    Geometry-Aware Implicit Memory for Video World Models

    arXiv:2606.02436v1 Announce Type: new Abstract: Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations leave its native context window. Explicit memories retain frames or online 3D recon…

  47. arXiv cs.CV TIER_1 English(EN) · Lei Zhu, Xing Cai, Yingjie Chen, Yiheng Li, Binxin Yang, Hao Liu, Jie Chen, Chen Li, Jing LYu ·

    OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation

    arXiv:2604.18326v2 Announce Type: replace Abstract: Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidelity human-centric videos in complex, real-world physical scenes remains a si…

  48. arXiv cs.CV TIER_1 English(EN) · Qixin Hu, Shuai Yang, Wei Huang, Song Han, Yukang Chen ·

    LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

    arXiv:2606.02553v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion enables variable-length synthesis, but long-horizon generation often suffers from accumulated errors and identity drift. For efficiency, existing methods commonly adopt sliding-window attention du…

  49. arXiv cs.CV TIER_1 English(EN) · Minseok Joo, Dogyun Park, Taehoon Lee, Kyujin Lee, Hyunwoo J. Kim ·

    Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation

    arXiv:2606.02479v1 Announce Type: new Abstract: Maintaining long-term geometric consistency remains challenging for long-horizon autoregressive video generation. Memory-augmented generative models address this by retrieving historical frames, but their effectiveness depends on tw…

  50. arXiv cs.CV TIER_1 English(EN) · Shengjun Zhang, Zhang Zhang, Simin Huang, Zhenyu Tang, Hanyang Wang, Chensheng Dai, Min Chen, Yifan Li, Yuxin Li, Yingjie Chen, Hao Liu, Chen Li, Yueqi Duan ·

    MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

    arXiv:2606.00793v1 Announce Type: new Abstract: Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fundamental gap persists between visually plausible video generation and the functio…

  51. arXiv cs.CV TIER_1 English(EN) · Yukang Chen ·

    LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

    Autoregressive (AR) video diffusion enables variable-length synthesis, but long-horizon generation often suffers from accumulated errors and identity drift. For efficiency, existing methods commonly adopt sliding-window attention during generation. This creates an irreversible ge…

  52. arXiv cs.CV TIER_1 English(EN) · Hyunwoo J. Kim ·

    Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation

    Maintaining long-term geometric consistency remains challenging for long-horizon autoregressive video generation. Memory-augmented generative models address this by retrieving historical frames, but their effectiveness depends on two key design choices: what 3D-geometric evidence…

  53. arXiv cs.CV TIER_1 English(EN) · Qi Fan ·

    Geometry-Aware Implicit Memory for Video World Models

    Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations leave its native context window. Explicit memories retain frames or online 3D reconstructions, which can suffer from heuristic retr…

  54. arXiv cs.CV TIER_1 English(EN) · Lin Zhao, Yushu Wu, Yifan Gong, Yanzhi Wang, Pu Zhao ·

    OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation

    arXiv:2605.30519v1 Announce Type: new Abstract: Autoregressive (AR) video generation extends videos by producing latent chunks sequentially, but scaling to long videos requires repeated access to a growing historical KV cache. Existing methods reduce this cost by truncating the K…

  55. dev.to — Claude Code tag TIER_1 English(EN) · Aliaksei Zelianouski ·

    My video generation pipeline that built itself

    <p>Let me show you something cool. This two-minute video was built by Claude Code from a single prompt.</p> <p> </p> <p>Okay — one prompt and about thirty follow-ups. And then twenty more after Claude Code fumbled a git command and wiped out half of my video-editing material (don…

  56. r/StableDiffusion TIER_2 English(EN) · /u/Sporeboss ·

    PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory (github link in description and 400gb training dataset)

    <!-- SC_OFF --><div class="md"><p><a href="https://ys-imtech.github.io/projects/PermaVid/">https://ys-imtech.github.io/projects/PermaVid/</a></p> <p><a href="https://huggingface.co/datasets/ysmikey/PermaVid_datasets">https://huggingface.co/datasets/ysmikey/PermaVid_datasets</a></…

  57. r/StableDiffusion TIER_2 Dansk(DA) · /u/hpyfox ·

    Finding old video generation models

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1u9ava9/finding_old_video_generation_models/"> <img alt="Finding old video generation models" src="https://preview.redd.it/zsu8inzlh28h1.gif?frame=1&amp;width=140&amp;height=140&amp;crop=1:1,smart&amp;aut…

  58. r/StableDiffusion TIER_2 English(EN) · /u/DesireForDopamine ·

    SCAIL-2 Infinity — a single node for unlimited-length video (no more chaining samplers) + Pusa LoRA integration

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1u7m255/scail2_infinity_a_single_node_for_unlimitedlength/"> <img alt="SCAIL-2 Infinity — a single node for unlimited-length video (no more chaining samplers) + Pusa LoRA integration" src="https://externa…

  59. r/StableDiffusion TIER_2 Dansk(DA) · /u/DoskvolDenizen ·

    Advice for overall image generation pipeline for video keyframes

    <!-- SC_OFF --><div class="md"><p>Been learning stable diffusion text to image and image to video with comfyui for a few months.</p> <p>Now I have so many tools at my disposal that I'm feeling a bit lost, so I'm hoping that people in here won't mind sharing some advice on an over…