Researchers have investigated the limitations of video world models in tracking unobserved states, even when they exhibit high visual fidelity. Using a "Shell Game" task, they found that models like Bidirectional and autoregressive Transformers, Mamba, and linear attention struggled with extrapolation beyond a certain number of swaps, indicating they did not maintain the hidden state of the world. The study suggests that the pixel-based diffusion target does not supervise the unseen hidden state, forcing models to re-derive arrangements from history. However, two mechanisms were identified that do extrapolate: linear attention with negative transition eigenvalues and a nonlinear fast weight mechanism in TTT, both of which carry and revise state across chunks. AI
IMPACT This research highlights critical limitations in current video world models, suggesting a need for architectures that can robustly track unobserved states for more reliable simulation and exploration.
RANK_REASON Research paper published on arXiv detailing findings about video world models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Bidirectional and autoregressive Transformers
- linear attention
- Mamba
- Shell Game
- video world models
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →