arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

几何声学模拟器复杂度对训练后的多声源定位器的影响

Influence of Geometrical Acoustic Simulator Complexity on a Trained Multisource Localizer

Fabian Staub, Nils Meyer-Kahlen, Thomas Deppisch, Sergio de las Heras, Florian Klein, Stephan Werner, Johannes M. Arend

arXiv 2610.05055首次发表:更新:

AI 中文总结

本研究探讨几何声学模拟复杂度对多声源定位模型真实性能的影响,发现中等复杂度镜像源模拟已足够,随机化镜像源位置是提升泛化的高效手段。

AI 中文摘要

基于学习的声学声源定位与检测(SSLD)需要覆盖多种声学条件的大型标注数据集。由于获取实测数据成本高昂,训练通常使用模拟数据,而实际设备必须在真实世界条件下运行。然而,更高的模拟复杂度会增加数据生成成本,且可靠泛化所需的复杂度仍不明确。在实践中,SSLD方法通常依赖高效的几何声学模拟,典型的是镜像源方法。本研究探讨了几何声学模拟复杂度如何影响常见多声源SSLD模型在真实世界中的性能。我们使用从消声条件到高阶镜像源模拟的多种模拟器训练模型,可选包括扩散混响、阵列模拟和随机化镜像源位置,并在三个基于测量的测试数据集上评估。结果表明,消声模拟不足,而中等复杂度的镜像源模拟已能提供强大的真实世界性能。进一步增加复杂度仅带来边际收益,最佳性能出现在使用最高测试阶数的镜像源、阵列模拟和随机化镜像源位置时。对于该配置,实测域性能接近域内模拟性能,表明进一步增加模拟复杂度的收益有限。这些发现展示了如何在未来的数据驱动SSLD中利用复杂度-性能权衡,并强调镜像源随机化是提升泛化能力的有效途径。

英文摘要

Learning-based acoustic sound source localization and detection (SSLD) requires large labeled datasets covering diverse acoustic conditions. Since obtaining measured data is costly, training commonly uses simulated data, while practical devices must operate under real-world conditions. However, higher simulation complexity increases data-generation cost, and the complexity required for reliable generalization remains unclear. In practice, SSLD methods often rely on efficient geometrical acoustic simulation, typically the image-source method. This work investigates how geometrical acoustic simulation complexity affects the real-world performance of a common multi-source SSLD model. We train the model with simulators ranging from anechoic conditions to high-order image-source simulations, optionally including diffuse reverberation, array simulation, and randomized image-source positions, and evaluate them on three measurement-based test datasets. Results show that anechoic simulation is insufficient, while medium-complexity image-source simulations already provide strong real-world performance. Further increases in complexity yield only marginal gains, with the best performance obtained using the highest tested image-source order, array simulation, and randomized image-source positions. For this configuration, measured-domain performance approaches within-domain simulated performance, suggesting limited benefit from further increasing simulation complexity. These findings show how the complexity-performance trade-off can be exploited in future data-driven SSLD, and highlight image-source randomization as an efficient way to improve generalization.

CommentsSubmitted to IEEE International Conference of Acoustics, Speech, and Signal Processing (IEEE ICASSP 2027)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑