arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于逐通道掩码的空间分布式麦克风在线掩码波束成形研究

A Study on Online Mask-based Beamforming Using Per-channel Masking for Spatially Distributed Microphones

Wiebke Middelberg, Svantje Voit, Simon Doclo, Ryan Corey

arXiv 2607.26623首次发表:更新:

AI 中文总结

该研究针对空间分布式麦克风的语音增强问题,提出逐通道掩码的多通道掩码波束成形方法,采用帧因果在线实现,实验表明其在麦克风信号差异显著时优于单一掩码,且鲁棒性良好。

AI 中文摘要

掩码波束成形是一种流行的与几何无关的语音增强方法,通常在所有麦克风上应用单一掩码以估计所需的协方差矩阵。该策略对于紧凑阵列有效,但对于空间分布式麦克风可能并非最优,因为不同麦克风间的信号特性可能存在显著差异。为有效捕捉麦克风间的空间多样性,我们将掩码波束成形扩展为多通道形式,在协方差估计前,每个麦克风由独立掩码预滤波。针对频谱-时间非平稳性导致的时变声学场景,我们采用带滑动窗口的帧因果在线实现。对模拟紧凑阵列与分布式麦克风的实验表明,当麦克风信号差异显著时,多通道掩码相比单一掩码更具优势,而在紧凑阵列中性能相近。我们还通过对比神谕理想比率掩码与基于DNN的盲掩码估计,验证了多通道掩码方法的鲁棒性。

英文摘要

Mask-based beamforming is a popular geometry-agnostic approach for speech enhancement, typically applying a single mask across all microphones to estimate the required covariance matrices. While effective for compact arrays, this strategy may be suboptimal for spatially distributed microphones, where signal characteristics may vary strongly across microphones. To effectively capture the spatial diversity across microphones, we extend the mask-based beamformer to a multi-channel formulation, where each microphone is pre-filtered by a separate mask before covariance estimation. To address time-varying acoustic scenes, caused by spectro-temporal nonstationarity, we adopt a frame-causal online implementation with a sliding window. Experiments with simulated compact arrays and distributed microphones show that multi-channel masking yields a benefit over using a single mask when microphone signals differ substantially, while retaining similar performance in compact arrays. We further demonstrate the robustness of the multi-channel masking approach by comparing oracle ideal ratio masks to blind DNN-based mask estimation.

CommentsAccepted for publication at IWAENC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑