arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17773cs.MM

FillGauss:用于3D高斯点云的细粒度填充感知撞击声生成

FillGauss: Fine-Grained Filling-Aware Impact Sound Generation for 3D Gaussian Splatting

Chen Yang, Ganye Wen, Bin Huang, Jiayi Lyu, Zehai Niu, Linlin Shen, Jinbao Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对多模态AI中从视觉合成撞击声的挑战,提出细粒度填充感知撞击声生成任务。基于FillImpact数据集,构建FillGauss框架,融合多种要素实现位置、撞击物及填充感知音频生成,达物理基础跨模态音频生成新水平。

中文摘要 AI 辅助

在多模态人工智能中,从视觉观察合成物理上合理的撞击声仍然是一个巨大挑战。现有3D感知音频生成方法主要对空心刚体的表面几何形状进行建模,忽略了内部填充状态这一关键物理因素。为此定义了细粒度填充感知撞击声生成任务。首先引入细粒度填充感知数据集(FillImpact),包含88个不同真实物体的5000多个声学记录。在此基础上提出FillGauss框架,将3D高斯点云与内部状态条件相结合用于声音生成。实验表明该方法能生成符合物理原理的高保真撞击声,为物理基础的跨模态音频生成建立了新的技术水平。

英文摘要

Synthesizing physically plausible impact sounds from visual observations remains a great challenge in multi-modal AI. Existing 3D-aware audio generation methods primarily model the surface geometry of hollow rigid bodies. However, they fundamentally overlook internal filling states, a critical physical factor that drastically modulates acoustic resonance and damping. To address this issue, we have defined a new task called Fine-Grained Filling-Aware Impact Sound Generation. As a foundational step, we first introduce the fine-grained fill-aware dataset (FillImpact), a pioneering multi-modal collection comprising over 5,000 rigorous acoustic recordings from 88 diverse real-world objects. It captures impact interactions with varying internal contents (i.e., water, rice), a continuous range of fill levels, and distinct striker materials. Furthermore, comprehensive acoustic analysis confirms that the collected data closely aligns with established physical laws governing acoustic resonance and damping, indicating its suitability for physically grounded modeling. Building on this dataset, we propose a novel generative framework (FillGauss) that integrates 3D Gaussian Splatting (3DGS) with internal state conditioning for sound generation. By fusing 3DGS geometric features, precise 3D spatial strike coordinates, and fine-grained textual physical conditions within a latent diffusion architecture, FillGauss enables position-aware, striker-aware, and filling-aware audio generation. Extensive experiments demonstrate that our approach could generate high-fidelity impact sounds that adhere to underlying physical principles, establishing a new state-of-the-art for physically grounded cross-modal audio generation.

↑