arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

几何信息感知的分布式声学场景理解

Geometry-Informed Distributed Acoustic Scene Understanding

Yiyuan Yang, Shitong Xu, Niki Trigoni, Andrew Markham

arXiv 2609.08026首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出几何信息感知的分布式声学场景理解框架,融合分布式麦克风、音频频谱图变换器与拓扑感知图神经网络,解码语义三元组并结合大语言模型,实现多房间场景的空间理解与物理一致叙述。

AI 中文摘要

多房间环境中的声学场景理解是一项困难的任务。现有大多数系统使用单个集中式麦克风阵列,且常常因墙壁和门阻挡声音信号而失效。为应对这一挑战,我们提出了一种几何信息感知的分布式声学场景理解框架。我们的系统利用分布式麦克风,并使用音频频谱图变换器和拓扑感知图神经网络来融合时空声学特征。随后,这些特征被解码为离散语义三元组。最后,一个冻结的大语言模型将这些符号观测与环境几何信息相结合。这使得系统能够执行空间理解、推断合理的缺失转换,并生成物理上一致的场景叙述。在自定义多房间模拟器上的实验表明,我们的框架优于集中式基线,并在模拟遮挡下提升了空间一致性。

英文摘要

Acoustic scene understanding in multi-room environments is a difficult task. Most existing systems use a single centralized microphone array, and they often fail because walls and doors block sound signals. To address this challenge, we propose a geometry-informed distributed acoustic scene understanding framework. Our system leverages distributed microphones and uses an audio spectrogram transformer and a topology-aware graph neural network to fuse spatio-temporal acoustic features. Then, these features are decoded into discrete semantic triplets. Finally, a frozen large language model combines these symbolic observations with the environmental geometry. This allows the system to perform spatial understanding, infer plausible missing transitions, and generate a physically consistent narrative of the scene. Experiments on a custom multi-room simulator demonstrate that our framework outperforms centralized baselines and improves spatial consistency under simulated occlusion.

CommentsAccepted by Interspeech 2026 Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑