发表机构
Microsoft; University of Michigan(微软; 密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对在线音视频目标说话人提取的听感质量与前瞻限制问题,提出AV-PVQE方法,通过个性化语音增强模型引入口型特征并联合微调,将目标混淆率从46%降至1.6%,延迟仅20毫秒,在合成和会议基准上均优于现有提取器。
AI 中文摘要
在线音视频目标说话人提取旨在去除竞争语音,同时保持语音质量并限制前瞻。现有提取器是在合成混合语音上构建和评估分离性能的,导致听感质量和会议行为基本未得到测试。我们提出了音视频个性化语音质量增强(AV-PVQE),从另一方向满足这些要求。我们从个性化语音增强模型出发,该模型能以高质量重建所请求的语音,但在干净注册的情况下,在46%的双说话人混合语音中仍会混淆目标。在其说话人条件输入中加入口型特征,并联合微调视觉和重建网络,可将该混淆率降至1.6%,且无需未来帧,算法延迟仅为20毫秒。与在线自回归音视频提取器相比,AV-PVQE在两个合成基准上取得分离增益,在录制的会议上增益更大,并且在说话人数多于微调混合语音的片段上保持优势。在两个会议语料库的个性化P.835听感测试中,AV-PVQE相对该提取器将整体质量提升了0.57和0.63 MOS,平均评分与起始模型相近。保留和拒绝测试表明,当不存在竞争语音时,它能保持目标语音完整;当目标缺失时,它能抑制竞争语音。
英文摘要
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.