Scene2Sound:面向3D高斯世界的听觉锚定声景生成
Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds
浏览论文内容
中文总结 AI 辅助
Scene2Sound是一个无训练框架,通过听觉锚定和高斯集匹配为3D高斯世界生成空间一致的声景,在保留音频质量的同时解决了现有方法空间不一致的问题,经实验和用户研究验证有效。
中文摘要 AI 辅助
3D高斯溅射(3DGS)可将捕获或生成的图像转化为用户可自由探索的逼真3D世界模拟,但这些世界仍处于无声状态。由于现有音频生成方法以单张图像或视角为条件,其生成的声音与该观测结果绑定,在听者移动时无法保持一致。我们提出为给定3DGS世界生成空间一致声景的任务,该任务通过听觉锚定识别世界中哪些物体应发出声音,并将每个声音锚定到持久的3D位置,同时推出Scene2Sound,这是一个基于该锚定的无训练框架。我们的流程仅从输入世界中选择共同覆盖场景的视角,利用视觉语言模型识别发声物体,并通过高斯集匹配将多视图检测关联为3D实例,高斯集匹配用于测量每个检测渲染的高斯集之间的重叠。每个声源随后接收由标准基于物体的音频引擎在任意听者姿态下实时空间化的生成音频。我们进一步提出两个空间一致性指标:一个测试渲染音频是否对听者移动做出一致响应,另一个测试所声称的声源是否得到其放置的保留视图的支持。在一组精心挑选的生成3DGS世界以及从真实世界360度捕获生成的3DGS场景上,Scene2Sound在保留强单视角基线音频质量的同时,在单视角和单全景管道无法保持空间一致的地方实现了空间一致性,且用户研究证实了其感知益处。项目页面:this https URL。
英文摘要
3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these worlds remain silent. Because existing audio generation methods condition on a single image or viewpoint, their sound is tied to that observation and cannot stay consistent while a listener moves. We introduce the task of generating a spatially consistent soundscape for a given 3DGS world through auditory grounding, identifying which objects in the world should emit sound and anchoring each to a persistent 3D position, and present Scene2Sound, a training-free framework built on this grounding. From the input world alone, our pipeline selects viewpoints that jointly cover the scene, identifies sound-emitting objects with a vision-language model, and associates the multi-view detections into 3D instances through Gaussian set matching, which measures the overlap between the Gaussian sets that render each detection. Each source then receives generated audio that a standard object-based audio engine spatializes in real time at arbitrary listener poses. We further propose two spatial-consistency metrics, one testing whether rendered audio responds consistently to listener motion and one testing whether the claimed sources are supported by views held out from their placement. On a curated set of generated 3DGS worlds and on 3DGS scenes generated from real-world 360-degree captures, Scene2Sound preserves the audio quality of strong per-viewpoint baselines while remaining spatially consistent where per-viewpoint and single-panorama pipelines do not, and a user study confirms the perceptual benefit. Project page: https://masaki-lmd.github.io/scene2sound/.
发表机构
- Hokkaido University(北海道大学)
机构由 AI 辅助整理,请以论文原文为准。