EquiSELD:等变声音事件定位与检测网络的高效训练
EquiSELD: Efficient training of equivariant sound event localization and detection networks
- Radboud University(拉德堡德大学)
- Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
EquiSELD提出一种利用FOA信号O(3)对称性的等变注意力网络,通过成对流处理和Multi-ACCDOA读出,以更低训练成本在模拟和真实声景中超越先前等变及非等变SELD网络。
AI中文摘要:
一阶环境立体声(FOA)信号具有精确的O(3)对称性:FOA信号的旋转或反射会改变声源的到达方向,同时保持声源本身不变。先前利用FOA的这种空间对称性来提高声音事件检测与定位(SELD)系统的效率和鲁棒性的尝试,要么仅通过基于旋转的数据增强来学习对称性的近似,要么依赖计算成本高昂的方法来整合等变性。此外,先前的工作仅专注于SO(3)等变性,使得将O(3)等变性纳入SELD任务的潜力尚不明确。为解决这些局限性,我们开发了EquiSELD。这种等变注意力网络将一阶环境立体声处理为O(3)不变标量与等变强度向量的成对流,通过Multi-ACCDOA读出产生不变的活动幅度和等变的到达方向(DOA)。为了比较O(3)与SO(3)等变性的影响,我们设计了一个匹配的仅SO(3)变体。EquiSELD在带有实测房间冲激响应(RIR)的模拟场景和真实世界声景录音上均优于先前的等变网络,且训练成本仅为后者的一小部分。此外,EquiSELD在模拟的真实世界声景上超越了相似规模的非等变SELD网络的性能,并在真实世界声景上取得了具有竞争力的性能。
英文摘要:
First-order Ambisonics (FOA) signals exhibit exact O(3) symmetry: The rotation or reflection of the FOA signal modifies the direction of arrival of the sound sources, while preserving the sound sources themselves. Prior attempts to utilize this spatial symmetry of FOA to improve the efficiency and robustness of sound event detection and localization (SELD) systems either learned only an approximation of the symmetry through rotation-based augmentation or relied on computationally expensive methods to integrate equivariance. Furthermore, prior work focused exclusively on SO(3) equivariance , leaving the potential of incorporating O(3) equivariance for SELD tasks unclear. To address these limitations, we developed EquiSELD. This equivariant attention network processes first-order Ambisonics as paired streams of O(3)-invariant scalars and equivariant intensity vectors, producing an invariant activity magnitude and an equivariant DOA with a Multi-ACCDOA readout. To compare the impact of O(3) versus SO(3)-equivariance, we designed a matched SO(3)-only variant. EquiSELD outperforms prior equivariant networks on both simulated scenes with measured RIRs and recordings of real-world sound scenes at a fraction of the training cost. EquiSELD additionally surpasses the performance of non-equivariant SELD networks of a similar size on the simulated real-world sound scenes and achieves competitive performance on the real-world sound scenes.