arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38444cs.CVcs.SD

可听世界模型:面向3D世界的空间感知声音生成

Audible World Models: Spatially Aware Sound Generation for 3D Worlds

Duowen Chen, Jinjin He, Gouthaman KV, Sandeep Bangalore Venkatesh, Bo Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出可听世界模型,一种无需训练的框架,通过语义分层、几何锚定和声学传播,为3D世界生成随听者位置变化的空间音频,显著提升空间一致性。

中文摘要 AI 辅助

文本和图像条件的世界生成器可以创建视觉丰富的3D环境,然而这些世界往往保持沉默,或者仅依赖于完全由文本或渲染视频合成的配乐。尽管此类音频能够传达应该听到的内容,但它缺乏对声源位置以及感知声音应如何随听者移动而变化的显式表示。我们引入了可听世界模型,这是一种无需训练的框架,将声音融入生成的世界状态中。从文本提示开始,我们的系统构建一个全景3D代理,将其分离为语义层,识别产生声音的前景对象和环境背景区域,并为每个声音标签合成干音频。然后,它将声源锚定到重建的几何结构上,并使用几何声学传播渲染依赖于听者的空间音频。通过显式关联语义、几何和声音传播,该框架在保持持久声源位置的同时,使渲染音频适应听者视角和运动的变化。在80个生成场景上的实验表明,与文本、视频和全景条件基线相比,该框架在空间一致性上取得了显著提升,同时保持了具有竞争力的语义对齐。基于VLM的评估和人工评估进一步表明,我们的配乐在视听一致性、空间合理性和运动依赖行为方面更受青睐。

英文摘要

Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement. We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion. Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio-visual consistency, spatial plausibility, and motion-dependent behavior.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)
  • Dolby Laboratories(杜比实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑