真实对话语音增强中多说话人提取的挑战
Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement
- University of Sheffield(谢菲尔德大学)
- South Westphalia University of Applied Sciences(南威斯特法伦应用科学大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对真实对话中多说话人提取面临的沉默过多和注册语音不匹配挑战,提出新损失函数提升STOI和分段SNR,并分析不匹配影响。
AI中文摘要:
目标说话人提取和多说话人提取是在存在其他说话人和/或噪声的情况下,从期望说话人或期望的多个说话人中提取语音的技术。用于此任务的神经网络方法通常使用模拟数据集进行训练和评估,这些数据集包含平衡数量的目标语音和与目标语音紧密匹配的说话人注册样本。然而,在真实的多方对话中,参与者沉默的时间往往比说话的时间更长,并且他们的注册语音样本可能与对话中的目标语音存在显著差异。这些因素可能影响这些技术在真实对话录音上的训练和评估。本工作提出了一种新的损失函数,有助于减轻训练中过多沉默的影响,将STOI从0.55提高到0.60,频率加权分段信噪比从4.35提高到5.12。此外,还探讨了注册语音与目标语音之间不匹配的影响。
英文摘要:
Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.