研究人员推出了EchoWM,一个全模态世界模型,能够响应连续导航生成同步的视频、声音、音乐和语音。该模型支持第一人称和第三人称视角,在没有特定控制器的情况下学习摄像机-角色动力学。EchoWM利用互补数据引擎和渐进式训练,然后进行自回归后训练以实现长时程生成,在基准测试中实现了强大的轨迹跟踪和高质量的视觉效果。 AI
影响 该模型通过与连续导航同步多种模态,推动了生成式媒体的发展,可能对互动娱乐和模拟产生影响。
排序理由 该集群描述了一篇关于新AI模型EchoWM的研究论文,该论文已在Hugging Face和arXiv等平台上发布。
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
- AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
- AlayaWorld: Long-Horizon and Playable Video World Generation
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- EchoWM
- Gotit.pub
- Hugging Face
- MiniWorld: Democratizing the Training of Video World Models from Scratch
- ScienceCast
- Semantic Scholar API
- Vorch-Omni: Multi-Task Orchestration of Sight and Sound
- Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
- Wonder: Video World Model Done Better
AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →