arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ISCSLP 2026 真实场景视听语音增强挑战赛

The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge: A Benchmark with Natural Mixtures and Degraded Video

Kai Li, Wenze Ren, Junjie Li, Cheng Yu, Peijun Yang, Chien-yu Huang, Haibin Wu, Szu-Wei Fu, Wen-Chin Huang, Hsin-Min Wang, Xiaolin Hu, Ming Li, Yu Tsao, DeLiang Wang

arXiv 2608.23759首次发表:更新:

发表机构

Real-World AVSE Challenge Organizing Committee(真实世界视听语音增强挑战赛组织委员会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究介绍ISCSLP 2026真实场景视听语音增强挑战赛,设置含真实混合、合成重混等赛道,公布基线结果与相关资源,以评估视听语音增强在自然场景下的性能。

AI 中文摘要

视听语音增强(AVSE)利用目标说话人的视觉语音线索,从噪声或重叠语音中恢复该说话人的语音。许多广泛使用的协议从单独录制的音频源构建混合信号,并假设视频可靠,导致其在自然重叠和视频失效情况下的性能未得到充分表征。本次真实场景视听语音增强挑战赛评估两个相关设置:赛道1包含两个场景,一是两名说话人同时录制的真实混合信号,无对应干净参考信号;二是通过手动混合两名说话人单独录制的语音获得的合成重混信号,有干净参考信号可用。赛道2复用音频,但将其与退化的目标视频配对,并包含额外的3米远场录制内容。开发集和测试集的说话人不重叠。评估指标包括干净波形保真度、学习质量估计、转录准确率和说话人识别。在开发集的重混任务中,基线模型在赛道1的SI-SDR为-4.069 dB、STOI为0.388,在赛道2的SI-SDR为-2.851 dB、STOI为0.470。我们发布了AV-ConvTasNet的检查点、离线评估器,以及开发集和测试集的官方基线结果。

英文摘要

Audio-visual speech enhancement (AVSE) uses a target speaker's visible articulation to recover that speaker's voice from overlapped speech. Yet most evaluation protocols rely on synthetic mixtures and reliable video. The ISCSLP 2026 Real-World AVSE Challenge addresses both gaps. Track~1 combines naturally recorded two-talker mixtures, which lack a clean reference, with reference-available remixes of the same speakers; Track~2 additionally degrades the target video in five ways and adds 3-m far-field recordings. Sixteen and twelve teams were ranked on speaker-disjoint test data by rank averaging over waveform fidelity, predicted quality, transcription accuracy, and speaker similarity. On Track~1 remixes, the best system reaches 12.7~dB SI-SDR and 0.85 STOI, but natural recordings remain harder: even the lowest CER rises from 9.0\% to 14.7\%. The leading systems use video mainly for speaker attribution rather than signal reconstruction, and the top two lose under 0.5~dB SI-SDR on Track~2. UTMOS and DNSMOS rank systems differently from the other metrics, so no single metric captures target-speech recovery. We release the baselines, evaluator, and official results.

CommentsThe First Real-World Audio-Visual Speech Enhancement (AVSE) Challenge; Submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑