PulseAugur
实时 10:11:35
English(EN) Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

Scene2Sound 为 3D 高斯世界生成一致的音频

研究人员推出了 Scene2Sound,一个为 3D 高斯世界生成一致声景的新框架。该系统通过创建空间连贯且在听众穿梭于模拟环境时保持稳定的音频,解决了现有音频生成方法的局限性。Scene2Sound 利用视觉语言模型识别 3D 世界中的发声物体,并采用高斯集合匹配将这些检测结果与持久的 3D 位置关联起来,从而实现实时音频空间化。 AI

影响 通过添加空间一致的音频,实现了更具沉浸感和交互性的 3D 环境。

排序理由 该集群包含一篇研究论文,详细介绍了在 3D 环境中生成音频的新方法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Scene2Sound 为 3D 高斯世界生成一致的音频

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Masaki Yoshida, Ren Togo, Takahiro Ogawa, Miki Haseyama ·

    Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

    arXiv:2608.00463v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these worlds remain silent. Because existing audio generation methods condition on a single…