English(EN)MiniWorld: Democratizing the Training of Video World Models from Scratch
新的交互式视频世界模型框架出现
作者PulseAugur 编辑部·[5 个来源]·
研究人员推出了两个新的视频世界模型框架,这对于具身AI和交互式模拟至关重要。第一个框架HelloWorld通过允许角色响应诸如转弯或挥手之类的提示,实现了用户与视频环境中的角色之间的社交互动。它利用了自蒸馏管道和新颖的推理模块来实现响应的时间局部化。第二个框架MiniWorld旨在通过提供一个轻量级、可复现的系统,该系统可以在适度的计算资源上从头开始训练,从而实现这些模型的训练民主化。MiniWorld采用了块因果视频扩散Transformer和流匹配,使其能够被更广泛的研究使用。
AI
Despite the remarkable recent progress of video world models, social interaction between users and the characters within these worlds remains unsupported. To fill this gap, we present HelloWorld, a video world model that enables social interaction with in-world characters. With a…
Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions. Unlike conventional video generation models that primarily capture visual appearance and motion, v…
arXiv:2608.05070v1 Announce Type: new Abstract: Despite the remarkable recent progress of video world models, social interaction between users and the characters within these worlds remains unsupported. To fill this gap, we present HelloWorld, a video world model that enables soc…
arXiv:2604.21686v2 Announce Type: replace Abstract: Unlike text- or image-driven video generation, an interactive world model is driven by actions: the user acts, and the world responds. Two obstacles stand in the way of fair and comprehensive evaluation. First, models take actio…
arXiv:2608.01127v1 Announce Type: new Abstract: Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions. Unlike conventional video generation models that p…