arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经视场用于可穿戴麦克风阵列的双耳信号匹配

Neural Field-of-View for Binaural Signal Matching with Wearable Microphone Arrays

Matan Yifrach, Boaz Rafaely

arXiv 2609.28343首次发表:更新:

发表机构

School of Electrical and Computer Engineering, Ben-Gurion University of the Negev(内盖夫本-古里安大学电气与计算机工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出FoV-BSM-Net,一种基于卷积循环神经网络端到端学习视场参数的双耳信号匹配方法,无需显式声源定位,在模拟房间中显著提升双耳信号质量,尤其在高直达-混响比条件下。

AI 中文摘要

空间音频在增强现实和虚拟现实等应用中的日益普及,推动了针对麦克风数量有限的可穿戴阵列的双耳重放方法的发展。双耳信号匹配(BSM)是其中一种方法,在漫射场假设下产生高质量的双耳信号,但在直达声占主导的高直达-混响比(DRR)条件下性能会下降。先前的扩展引入了视场(FoV)加权,要么采用固定孔径,要么基于显式声源定位,但这些方法受限于粗糙的空间覆盖或对定位估计精度的依赖。本文提出了FoV-BSM-Net,一种信号相关的FoV-BSM公式,通过使用卷积循环神经网络从麦克风信号中端到端学习FoV参数,避免了显式声源估计。该方法在不同混响条件下的模拟房间中进行了评估,并与BSM和固定FoV-BSM基线进行了比较。结果表明,FoV-BSM-Net在双耳NMSE和耳间线索误差方面持续优于BSM,且增益随DRR增大而增加,同时感知评估进一步支持了其在低和高DRR条件下均显著优于两个基线的优势。

英文摘要

The growing use of spatial audio in applications such as augmented and virtual reality has driven the development of binaural reproduction methods for wearable arrays with a limited number of microphones. Binaural signal matching (BSM) is one such method, producing high-quality binaural signals under a diffuse-field assumption, but degrading at high direct-to-reverberant ratios (DRR) where the direct sound dominates. Previous extensions incorporate Field-of-View (FoV) weighting, either with fixed apertures or based on explicit source localization, but these approaches are limited by coarse spatial coverage or reliance on localization estimation accuracy. This paper introduces FoV-BSM-Net, a signal-dependent FoV-BSM formulation that avoids explicit source estimation by learning the FoV parameters end-to-end from the microphone signals using a Convolutional Recurrent Neural Network. The method is evaluated in simulated rooms across varying reverberation conditions, and compared against BSM and a fixed FoV-BSM baseline. Results show that FoV-BSM-Net consistently improves over BSM, with gains that grow with DRR in both binaural NMSE and interaural cue errors, and are further supported by perceptual evaluation showing a substantial advantage over both baselines across low and high DRR conditions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑