arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

注意麦克风差距:潜在声学映射的阵列上采样策略基准测试

Mind the Microphone Gap: Benchmarking Array Upsampling Strategies for Latent Acoustic Mapping

Philipp Schmidt, Huw Cheston, Juan Azcarreta, Adrian Stepien, Çağdaş Bilen, Iran R. Roman

arXiv 2607.24463首次发表:更新:

AI 中文总结

研究潜在声学映射在稀疏4通道阵列中退化问题,对多种上采样架构进行基准测试,包括卷积网络等。通过不同训练方式研究与LAM对齐效果,发现原始全分辨率LAM最强,单独训练的轻量级模型最具竞争力,且表示对齐比模型复杂性更重要。

AI 中文摘要

潜在声学映射(LAM)是一种自监督学习方法,可从多通道录音中生成高分辨率球形声学图,且无需标记数据,在到达方向基准测试中与监督基线相匹配。然而,LAM在稀疏4通道阵列中会显著退化,因为低分辨率互谱矩阵捕获的空间信息远少于LAM设计所针对的32通道输入。我们对多种上采样架构进行了基准测试,包括轻量级卷积网络、迭代反投影模型、物理信息网络和生成对抗方法。我们还研究了通过联合训练或在不同阶段训练这些上采样器与LAM对齐是否有助于保留LAM所依赖的空间结构。结果表明,原始的全分辨率LAM最强,单独训练的轻量级模型是最具竞争力的学习方法,并且上采样器与LAM之间的表示对齐比模型复杂性更重要。

英文摘要

Latent Acoustic Mapping (LAM) is a self-supervised learning method that generates high-resolution spherical acoustic maps from multichannel recordings without labelled data, matching supervised baselines on direction-of-arrival benchmarks. However, LAM degrades significantly with sparse 4-channel arrays, as the low-resolution cross-spectral matrix captures far less spatial information than the 32-channel inputs LAM was designed for. We benchmark a diverse set of upsampling architectures, spanning lightweight convolutional networks, iterative back-projection models, physics-informed networks, and generative adversarial approaches. We also study whether aligning these upsamplers with LAM by training them jointly or in different stages helps preserve the spatial structure that LAM depends on. Results show that the original full-resolution LAM is the strongest, that separately trained lightweight models are the most competitive learned approaches, and that representation alignment between the upsampler and LAM matters more than model complexity.

CommentsIWAENC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑