AI 中文总结
该技术报告介绍MERL在实时语音分离挑战中的提交内容,基于基线模型,聚焦数据准备与清理,分四阶段训练系统。提交获第一名,凸显高质量数据准备的重要性,还通过对抗攻击研究了DNSMOS和说话者相似度的鲁棒性。
AI 中文摘要
目标语音提取(TSE)很大程度上由基于神经网络的方法主导,这些方法在合成的完全重叠数据上进行训练和评估。实时语音分离挑战旨在提高在真实世界远场噪声和混响录音上的性能。本技术报告描述了MERL提交给该挑战的内容。我们基于基线模型,主要关注数据准备和清理。系统分四个阶段训练,先在完全重叠混合数据上预训练,再用带伪目标的真实远场噪声录音调整模型。我们的提交在第二赛道获第一名,证明高质量数据准备的关键重要性。此外,我们发现DNSMOS和说话者相似度易受过度优化影响,通过对抗攻击研究其鲁棒性,结果表明二者可被驱动到极端值而不降低令牌错误率或基于VAD的F1分数。
英文摘要
Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims to advance performance on real-world far-field noisy and reverberant recordings. This technical report describes MERL's submission to the Real-TSE Challenge. Rather than proposing a novel model architecture, we built upon the baseline model and focused primarily on data preparation and cleaning. Our system was trained in four stages, beginning with pre-training on fully overlapped mixtures and simulated multi-talker conversations with noise and reverberation applied to both the mixture and the enrollment utterances. We then adapted the model to real-world conditions using noisy far-field recordings with pseudo-targets derived from processed close-talk microphone signals. Our submission achieved first place in the second track, demonstrating the critical importance of high-quality data preparation. Furthermore, we observed that DNSMOS and speaker similarity are susceptible to over-optimization, motivating an investigation of their robustness using adversarial attacks. The results show that both metrics can be driven to extreme values without degrading the token error rate or the VAD-based F1 score.