arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18040eess.AS

面向不规则麦克风阵列SELD的任务导向神经FOA编码

Task-oriented neural FOA encoding for SELD from irregular microphone arrays

Jiachen Liu, Yin Cao, Ming Wu, Jun Yang

AI总结:

针对不规则麦克风阵列,提出两阶段SELD框架,通过神经残差编码器和教师-学生知识蒸馏学习任务导向FOA表示,实验证明可提升SELD性能并降低定位误差。

AI中文摘要:

声音事件定位与检测(SELD)系统通常依赖一阶环绕声(FOA)输入,而从不规则麦克风阵列获取有用的FOA表示仍然具有挑战性。本文提出了一种两阶段SELD框架,该框架从麦克风阵列信号中学习任务导向的、与FOA兼容的表示。首先,神经残差编码器通过信号相关的校正来细化传统的FOA编码。然后,教师-学生方案通过帧级置换不变知识蒸馏,从理论FOA表示中转移事件和空间知识。在具有四面体和12通道Benchmark阵列的合成场景以及来自LOCATA数据集的真实固定声源录音上的实验表明,教师指导持续改善下游SELD性能,并显著降低定位误差。信号级分析进一步表明,较低的FOA重建误差不一定对应更好的SELD性能,这表明蒸馏表示主要针对任务相关的空间信息进行优化,而非严格的FOA重建。

英文摘要:

Sound event localization and detection (SELD) systems often rely on first-order Ambisonics (FOA) input, whereas obtaining useful FOA representations from irregular microphone arrays remains challenging. This paper proposes a two-stage SELD framework that learns a task-oriented, FOA-compatible representation from microphone-array signals. A neural residual encoder first refines conventional FOA encoding through a signal-dependent correction. A teacher--student scheme then transfers event and spatial knowledge from theoretical FOA representations through frame-level permutation-invariant knowledge distillation. Experiments on synthetic scenes with tetrahedral and 12-channel Benchmark arrays, together with real stationary-source recordings from the LOCATA dataset, show that teacher guidance consistently improves downstream SELD performance and substantially reduces localization error. Signal-level analysis further shows that lower FOA reconstruction error does not necessarily correspond to better SELD performance, indicating that the distilled representation is optimized primarily for task-relevant spatial information rather than strict FOA reconstruction.

补充信息

↑