发表机构
University of Maryland, College Park(马里兰大学帕克分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对现有碰撞声建模方法成本高或需大量数据的问题,提出AV-MSF表示,结合3D高斯溅射与模态参数实现少样本重建,在真实数据集上达最优渲染性能,还支持接触定位等下游应用。
AI 中文摘要
尽管现代三维重建在建模物体几何与外观方面表现出色,但很大程度上忽略了物理交互所揭示的丰富声学线索。物体碰撞声传递出材料、刚度和结构特性,这些特性可与视觉信息形成互补。然而,现有的碰撞声建模方法要么依赖成本高昂的基于物理的仿真,要么需要大量数据集才能以纯数据驱动的方式实现泛化。本文提出一种新颖的物体级声学表示——视听模态声场(AV-MSF),该表示从多视图图像和仅少量碰撞声录音中重建得到。AV-MSF基于三维高斯溅射(3D Gaussian Splatting)并结合密集三维视觉特征构建,以提供强几何感知先验,同时用紧凑且具有物理意义的模态参数表示碰撞声场,从而实现稳健的少样本重建。在两个真实世界数据集上开展的实验表明,AV-MSF在碰撞声渲染方面达到了当前最优性能,优于基于物理的基线方法和数据驱动的基线方法。此外,本文还展示了该表示支持的下游应用,包括接触定位和物体声音编辑。
英文摘要
While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling approaches either rely on expensive physics-based simulation or require large datasets to generalize in a purely data-driven manner. We introduce Audio-Visual Modal Sound Field (AV-MSF), a novel object-level acoustic representation reconstructed from multi-view images and only a few impact sound recordings. AV-MSF builds on 3D Gaussian Splatting integrated with dense 3D visual feature to provide a strong geometry-aware prior, and represents the impact sound field using compact, physically meaningful modal parameters, enabling robust few-shot reconstruction. Experiments on two real-world datasets show that AV-MSF achieves state-of-the-art impact sound rendering, outperforming both physics-based and data-driven baselines. Furthermore, we demonstrate downstream applications enabled by our representation, including contact localization and object sound editing.
CommentsECCV 2026, Project page: https://zisenshao.github.io/AV-MSF/