PulseAugur
实时 11:11:35
English(EN) Latent Spatial Memory for Video World Models

新基准和框架推动视频世界建模发展

研究人员推出了“ImageTime”,这是一个旨在评估图像生成模型理解和表示时间变化能力的新基准。该基准通过要求模型生成一个动作的四个有序关键状态来评估时空一致性,超越了单一图像质量指标。此外,还开发了一个名为BiWM的新框架,利用双向自回归来推进开源交互式视频世界模型,旨在提高生成质量和推理速度。另一篇论文提出了一种用于视频世界模型的“潜在空间记忆”,将场景信息直接存储在扩散潜在空间中,从而显著加快生成速度并减小内存占用。 AI

影响 视频世界建模基准和框架的进步可能会加速生成式AI在视频和模拟领域的进展。

排序理由 多篇研究论文介绍了用于视频世界模型的新基准和框架。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 13 个来源。 我们如何撰写摘要 →

新基准和框架推动视频世界建模发展

报道来源 [13]

  1. arXiv cs.AI TIER_1 English(EN) · Abdulaziz Alyahya, Abdallah Al Siyabi, Markus R. Ernst, Luke Yang, Levin Kuhlmann, Gideon Kowadlo ·

    ARROW:用于鲁棒世界模型的增强回放

    arXiv:2603.11395v3 Announce Type: replace-cross Abstract: Continual reinforcement learning challenges agents to acquire new skills while retaining previously learned ones with the goal of improving performance in both past and future tasks. Most existing approaches rely on model-…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    MoVerse:全景高斯支架的实时视频世界建模

    We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the input observes only a small fraction of the environment, while interactive roaming requires a complete…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    MoVerse:全景高斯支架的实时视频世界建模

    MoVerse generates real-time interactive video from single images by creating 360° panoramas and 3D Gaussian scaffolds, enabling efficient rendering through diffusion-based techniques.

  4. arXiv cs.AI TIER_1 English(EN) · Xinrui Wu, Lichen Huang ·

    图像模型能想象时间吗?ImageTime:一个通过时空一致性探测视觉世界模型的新基准

    arXiv:2606.10620v1 Announce Type: cross Abstract: Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly understood. Practical workflows such as storyboarding, step-by-step illustration, r…

  5. arXiv cs.AI TIER_1 English(EN) · Shaohao Rui, Xiaofeng Mao, Zhanyu Zhang, Peijia Lin, Yansong Zhu, Yibo Zhang, Haibin Wan, Weijie Ma ·

    BiWM:通过双向自回归推进开源交互式视频世界模型

    arXiv:2606.10135v1 Announce Type: cross Abstract: Transitioning bidirectional video diffusion models into an autoregressive paradigm improves the interactivity of video world models, but existing causal pipelines need many stages (control fine-tuning, autoregressive training, cau…

  6. arXiv cs.AI TIER_1 English(EN) · Lichen Huang ·

    图像模型能想象时间吗?ImageTime:一个通过时空一致性探测视觉世界模型的新型基准

    Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly understood. Practical workflows such as storyboarding, step-by-step illustration, reference-guided editing, and video previsualizatio…

  7. arXiv cs.AI TIER_1 English(EN) · Jewon Yeom, Hanseul Kim, Jeongjae Park, Sungmok Jung, Jaejin Lee, Taesup Kim ·

    是什么让视频世界模型潜在表征与动作相关:预测而非重建

    arXiv:2606.07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces. We study this question through a unif…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    用于视频世界模型的潜在空间记忆

    Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computationally expensive, requiring repeated rendering and VAE encoding, and inherently lossy, as the round…

  9. Hugging Face Daily Papers TIER_1 English(EN) ·

    用于视频世界模型的潜在空间记忆

    Latent spatial memory for video world models stores 3D scene information directly in diffusion latent space, eliminating pixel-space reconstruction overhead and achieving faster generation with reduced memory usage.

  10. arXiv cs.CV TIER_1 English(EN) · Yang Zhou, Ziheng Wang, Yuqin Lu, Haofeng Liu, Jun Liang, Shengfeng He, Jing Li ·

    MoVerse:具有全景高斯支架的实时视频世界建模

    arXiv:2606.13376v1 Announce Type: new Abstract: We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the input observes only a small fraction of the environmen…

  11. arXiv cs.CV TIER_1 English(EN) · Jing Li ·

    MoVerse:全景高斯支架的实时视频世界建模

    We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the input observes only a small fraction of the environment, while interactive roaming requires a complete…

  12. arXiv cs.CV TIER_1 English(EN) · Weijie Wang, Haoyu Zhao, Yifan Yang, Feng Chen, Zeyu Zhang, Yefei He, Zicheng Duan, Donny Y. Chen, Yuqing Yang, Bohan Zhuang ·

    用于视频世界模型的潜在空间记忆

    arXiv:2606.09828v1 Announce Type: new Abstract: Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computationally expensive, requiring repeated rendering and …

  13. arXiv cs.CV TIER_1 English(EN) · Bohan Zhuang ·

    用于视频世界模型的潜在空间记忆

    Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computationally expensive, requiring repeated rendering and VAE encoding, and inherently lossy, as the round…