AI 中文总结
针对真实场景视听语音增强性能下降问题,提出“分离-然后关联”两阶段方法,经挑战赛数据集验证有效。
AI 中文摘要
视听语音增强(AVSE)旨在利用视觉线索从多说话人混合音频中提取目标语音。尽管近期研究在模拟数据集上取得了优异性能,但将其应用于真实场景视听录音时,性能往往会大幅下降。为缩小这一差距,ISCSLP 2026会议举办的真实场景AVSE挑战赛呼吁参与者设计适用于真实场景的AVSE实用解决方案,该场景中自然共存说话人重叠、声学干扰、房间混响和视觉退化问题。在我们提交给该挑战赛的方案中,提出了解耦的“分离-然后关联”方法,包含两个阶段:分离阶段使用仅音频训练的模型(即不使用视觉线索)将输入的多说话人混合音频分离为各个说话人信号;关联阶段使用视听CLIP模型,通过跨模态相似度匹配识别与目标说话人面部视频相似度最高的分离语音信号。在挑战赛数据集上的评估结果表明了所提方法的有效性。
英文摘要
Audio-visual speech enhancement (AVSE) aims at extracting target speech from multi-speaker mixtures by exploiting visual cues. Although recent studies have reported strong performance on simulated datasets, the performance, however, often drops dramatically when they are applied to real-world audio-visual recordings. To bridge this gap, the Real-World AVSE Challenge held in the ISCSLP 2026 conference calls for participants to design a practical solution for AVSE under real-world conditions, where speaker overlap, acoustic interferences, room reverberation and visual degradations naturally co-exist. In our submission to the challenge, we propose a decoupled separation-then-association approach. It consists of two stages: a separation stage in which a trained, audio-only model (i.e., not using visual cues) is used to separate input multi-speaker mixture to individual speaker signals, followed by an association stage, where an audio-visual CLIP model is used to identify the separated speech signal with the highest similarity with the target speaker's facial video via cross-modal similarity matching. Evaluation results on the challenge dataset show the effectiveness of our proposed approach.