发表机构
Duke Kunshan University; Johns Hopkins University; Wuhan University; The Chinese University of Hong Kong, Shenzhen(昆山杜克大学; 约翰斯·霍普金斯大学; 武汉大学; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对噪声下耳语转正常语音内容线索丢失问题,提出WhisperVC-AV,通过融合时间声学与唇部特征恢复内容,在多个ASR系统及噪声条件下显著降低CER,并保持说话人相似性。
AI 中文摘要
环境噪声掩盖了耳语到正常语音转换所需的内容线索。我们提出了WhisperVC-AV,它在保持原始WhisperVC转换模块不变的同时恢复声学内容特征。其上下文引导的恢复模块通过注意力机制和门控残差校正,将时间声学上下文与同步的唇部特征相结合。在AISHELL6-Whisper上的实验表明,与WhisperVC相比,在三个ASR系统上,对于干净语音以及所有六种信噪比(SNR)条件下的MUSAN噪声,字符错误率(CER)均更低。最大的改进出现在0 dB SNR下,其中Qwen3-ASR的CER从36.29%降至28.88%。WhisperVC-AV还提高了预测语音的质量,同时保持了说话人相似性。CER的改进在无需重新训练的情况下扩展到未见过的背景噪声,而视觉控制支持使用话语特定的唇部线索。音频示例可在我们的演示页面上获取。
英文摘要
Environmental noise obscures the content cues needed for whisper-to-normal conversion. We propose WhisperVC-AV, which restores acoustic content features while leaving the original WhisperVC conversion module unchanged. Its context-guided restoration module combines temporal acoustic context with synchronized lip features through attention and gated residual correction. Experiments on AISHELL6-Whisper show lower character error rates (CERs) than WhisperVC across three ASR systems, on clean speech and under all six signal-to-noise ratio (SNR) conditions with MUSAN noise. The largest gains occur at 0 dB SNR, where Qwen3-ASR CER falls from 36.29% to 28.88%. WhisperVC-AV also improves predicted speech quality while maintaining speaker similarity. The CER gains extend to unseen background noise without retraining, while visual controls support the use of utterance-specific lip cues. Audio examples are available on our demo page.
Comments5 pages, 2 figures. Submitted to ICASSP 2027. Audio demos: https://larry-ziyue-yin.github.io/demo-whispervc-av/