arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PG-SELD:物理引导的声音事件定位与检测

PG-SELD: Physics-Guided Sound Event Localization and Detection

Elad Cohen, Elad Dror Cohen, Arnon Netzer, Hai Victor Habi

arXiv 2609.39216首次发表:更新:

AI 中文总结

针对SELD跨环境泛化受限问题,提出物理引导的PG-SELD训练框架,通过自由场教师知识蒸馏对齐中间表示,在STARSS23上提升多种基线模型的泛化性能。

AI 中文摘要

声音事件定位与检测(SELD)旨在从多声道音频中联合识别声音事件并估计其到达方向。尽管最近的深度学习方法取得了强劲的性能,但它们在声学环境间的泛化能力仍然有限,因为房间混响将环境特有的特征引入了学习到的表示中。在本工作中,我们通过利用物理自由场模型作为与房间无关的参考来解决这一挑战。具体而言,我们提出了PG-SELD,一种将自由场与物理引导的知识蒸馏相结合的训练框架。我们的方法将混响信号中提取的中间表示与匹配声学场景下自由场教师产生的表示对齐。这种引导鼓励模型保留事件和定位相关信息,同时降低对房间特有特征的敏感性。在STARSS23基准上的实验结果表明,PG-SELD持续提升了多个基线SELD架构的泛化性能。

英文摘要

Sound event localization and detection (SELD) aims to jointly recognize sound events and estimate their directions of arrival from multichannel audio. Although recent deep learning approaches have achieved strong performance, their ability to generalize across acoustic environments remains limited, as room reverberation introduces environment-specific characteristics into the learned representations. In this work, we address this challenge by leveraging a physical free-field model as a room-independent reference. Specifically, we propose PG-SELD, a training framework that combines free-field with physics-guided knowledge distillation. Our approach aligns intermediate representations extracted from reverberant signals with those produced by a free-field teacher for matched acoustic scenes. This guidance encourages the model to preserve event- and localization-relevant information while reducing sensitivity to room-specific characteristics. Experimental results on the STARSS23 benchmark show that PG-SELD consistently improves the generalization performance of multiple baseline SELD architectures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑