京东已开源其交互式视听世界模型JoyAI-EchoWM。该模型能够生成同步的视频以及声音、音乐和语音,并实时响应用户导航。它在WBench导航基准测试中取得了约81.6分。 AI
影响 此次发布有助于多模态AI领域的增长,能够生成更同步、更具交互性的视听内容。
排序理由 特定AI模型的开源发布。
在 Mastodon — mastodon.social 阅读 →
AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →
京东已开源其交互式视听世界模型JoyAI-EchoWM。该模型能够生成同步的视频以及声音、音乐和语音,并实时响应用户导航。它在WBench导航基准测试中取得了约81.6分。 AI
影响 此次发布有助于多模态AI领域的增长,能够生成更同步、更具交互性的视听内容。
排序理由 特定AI模型的开源发布。
在 Mastodon — mastodon.social 阅读 →
AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →
完整方法见我们的编辑标准。
JD.com open-sourced JoyAI-EchoWM and EchoWM-Flash at JDD, scoring about 81.6–81.7 on WBench Navigation with unified camera-intent control and native audio-video generation, alongside plans for a large domestic GPU cluster and logistics Superbrain 3.0.
JD.com has open-sourced JoyAI-EchoWM, an interactive audiovisual world model that generates synchronised video with ambient sound, music and speech while responding in real time to user navigation. The model scored about 81.6 on WBench Navigation. https:// pandaily.com/jd-joyai-e…