arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UltraM2M:利用文本转录和混合约束进行弱监督语音增强

UltraM2M: Leveraging Text Transcripts and Mixture Constraints for Weakly-Supervised Speech Enhancement

Liu, Jiachen, Wu, Fulin, Wang, Zhong-Qiu

arXiv 2610.05155首次发表:更新:

AI 中文总结

UltraM2M在无监督M2M算法基础上引入文本转录和ASR损失作为弱监督,以增强混合约束的噪声抑制能力,在CHiME-4上验证有效。

AI 中文摘要

我们提出了UltraM2M,一种弱监督语音增强算法,它建立在最近的无监督混合到混合(M2M)算法之上。M2M通过在一组真实记录的带噪混响多通道混合信号上训练深度神经网络来估计目标语音和非目标信号,从而实现无监督语音增强。两个估计信号由所谓的混合约束(MC)损失进行惩罚,该损失约束它们重建观测到的混合信号。尽管已被证明有效,但混合约束可能过于薄弱,无法实现足够的噪声抑制。为了解决这个问题,UltraM2M扩展了M2M,进一步利用真实记录混合信号的文本转录来设计自动语音识别(ASR)损失,以惩罚估计的语音信号。ASR损失可以被视为一种弱监督形式,有助于无监督增强。在CHiME-4数据集上的评估结果显示了UltraM2M的有效性。

英文摘要

We propose UltraM2M, a weakly-supervised speech enhancement algorithm building upon the recent unsupervised mixture-to-mixture (M2M) algorithm. M2M realizes unsupervised speech enhancement by training deep neural networks on a set of real-recorded noisy-reverberant multi-channel mixture signals to estimate target speech and non-target signals. The two estimated signals are penalized by a so-called mixture-constraint (MC) loss, which constrains them to reconstruct the observed mixture signals. Although shown to be effective, the mixture constraint may be too weak to enable sufficient noise reduction. To deal with this, UltraM2M extends M2M by further leveraging text transcripts of real-recorded mixtures to design an automatic speech recognition (ASR) loss to penalize the estimated speech signal. The ASR loss can be viewed as a form of weak supervision that could help unsupervised enhancement. Evaluation results on the CHiME-4 dataset show the effectiveness of UltraM2M.

Commentsin submission

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑