发表机构
Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对麦克风数量限制ArrayDPS性能的问题,提出VM-ArrayDPS,通过添加高信噪比虚拟麦克风提供额外多通道一致性约束,显著提升2/3说话人盲语音分离性能。
AI 中文摘要
盲源分离(BSS)是信号处理中的一个基本问题,旨在无需关于源信号或混合过程的先验知识,即可从混合信号中分离出多个源信号。传统方法,如独立向量分析(IVA),利用源信号的统计独立性。近年来,基于扩散的方法通过利用强大的生成先验,已成为一种有前景的替代方案。其中,ArrayDPS将BSS问题表述为后验采样问题,并利用预训练的语音扩散模型来引导干净源信号的恢复。其分离能力的一个关键因素是多通道一致性(MC)目标,该目标强制估计的源信号通过估计的声学传递函数重建观测到的麦克风混合信号。然而,阵列中的麦克风数量通常有限,这限制了ArrayDPS的性能。为解决此问题,我们提出了VM-ArrayDPS,一种新颖的方法,通过添加具有更高信噪比(SNR)的虚拟麦克风来增强麦克风阵列,这些麦克风可以提供额外的MC约束以提升分离性能。实验结果表明,VM-ArrayDPS在2说话人和3说话人数据集上均显著优于ArrayDPS,展示了虚拟麦克风增强在提升BSS性能方面的有效性。我们还进行了消融研究,以展示虚拟麦克风数量及其带来的MC目标权重的影响。
英文摘要
Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process. Traditional approaches, such as Independent Vector Analysis (IVA) exploits statistical independence of sources. Recently, diffusion-based approaches have emerged as a promising alternative by leveraging powerful generative priors. Among them, ArrayDPS formulates BSS problem as a posterior sampling problem, and utilizes a pretrained speech diffusion model to guide the recovery of clean source signals. A key factor behind its separation capability is the multi-channel consistency (MC) objective, which enforces the estimated source signals to reconstruct the observed microphone mixtures through the estimated acoustic transfer functions. However, the number of microphones in the array is often limited, which constrains the performance of ArrayDPS. To address this issue, we propose VM-ArrayDPS, a novel method that augments the microphone array with virtual microphones with higher-SNR, these microphones can offer extra MC constraints to enhance the separation performance. Experimental results demonstrate that VM-ArrayDPS significantly outperforms ArrayDPS on both 2-speaker and 3-speaker datasets, showcasing the effectiveness of virtual microphone augmentation in improving BSS performance. We also did ablation studies to show the influence of the number of virtual microphones and weight of the MC objective brought by virtual microphones.
CommentsSubmitted to ICASSP 2027 and currently under review