PulseAugur
EN
LIVE 10:44:33
中文(ZH) 终于!世界模型进入“有声时代”:24FPS画面+48kHz立体声实时生成

HelixWorld 1.0: Real-time interactive audio-visual world model released

Researchers from Noiz AI, in collaboration with institutions like Hong Kong University of Science and Technology and Tsinghua University, have introduced HelixWorld 1.0, a novel interactive audio-visual world model. This model generates both visuals at 24 FPS and stereo audio at 48kHz in real-time, responding dynamically to user actions. Unlike previous models that generated static videos or added audio post-hoc, HelixWorld's Transformer architecture unifies audio and visual generation from the outset, ensuring synchronized and contextually relevant sensory experiences. The team plans to release the model weights and code openly in the coming weeks. AI

IMPACT Sets a new standard for immersive AI experiences by synchronizing real-time audio and visual generation, potentially accelerating development in interactive virtual environments.

RANK_REASON This is a release of a novel interactive audio-visual world model by a research team, with plans for open-sourcing. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on 量子位 (QbitAI) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

HelixWorld 1.0: Real-time interactive audio-visual world model released

COVERAGE [1]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 思邈 ·

    Finally! World Models Enter the 'Audio Era': 24FPS Video + 48kHz Stereo Real-time Generation

    即将完全开源