发表机构
Shenzhen Loop Area Institute; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen; The Chinese University of Hong Kong(深圳河套学院; 香港中文大学(深圳)人工智能学院; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对真实对话中轮流发言机制被忽视的问题,首次提出在线音视频目标说话人提取基准,并设计基于大语言模型的目标说话人语音活动投影模块,从重叠混合语音预测未来活动以引导低延迟分离器,结合历史、同步与预测性上下文,在多种骨干网络上持续提升流式提取性能,真实对话中获近1分贝增益。
AI 中文摘要
在面对面实时交流中,说话主体必须通过自然的停顿、轮流发言和反馈语来追踪目标说话人,且常常处于背景交叉对话的干扰之中。然而,大多数目标说话人提取(TSE)研究依赖于完全重叠或部分重叠的模拟混合语音,忽略了真实对话中的轮流发言机制。为此,我们首次引入了在线音视频目标说话人提取(AV-TSE)基准,该基准基于完整的二元交互,并带有独立的第三方干扰。观察到从语义、声学和面部线索中预测即将到来的活动有助于在线AV-TSE,我们提出了一种基于大语言模型(LLM)的目标说话人语音活动投影(TS-VAP)模块。与传统的分离说话人通道的VAP不同,该模块直接从重叠的混合语音中预测目标说话人和对话伙伴的未来活动,利用语音大语言模型的语言和对话知识,并利用这一预测来引导低延迟的分离器。我们进一步将这种预测性上下文与历史性和同步性说话人上下文相结合。实验表明,TS-VAP在多种AV-TSE骨干网络上持续提升了流式提取性能,并且进一步结合历史、同步和预测性上下文,在真实AV对话中获得了近1分贝的增益。项目页面:此HTTPS URL。
英文摘要
In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of real conversations. We therefore introduce, to our knowledge, the first benchmark for online audio-visual TSE (AV-TSE), built from intact dyadic interactions with independent third-party interference. Observing that anticipating upcoming activity from semantic, acoustic, and facial cues benefits online AV-TSE, we propose an LLM-based target-speaker voice activity projection (TS-VAP) module. Unlike conventional VAP with separated speaker channels, it forecasts the future activity of the target and conversational partner directly from the overlapping mixture, drawing on the linguistic and conversational knowledge of a speech-LLM, and uses this prediction to guide a low-latency separator. We further combine this predictive context with historical and synchronous speaker context. Experiments show that TS-VAP consistently improves streaming extraction across multiple AV-TSE backbones, and that further combining historical, synchronous, and predictive context yields nearly 1 dB gain on real AV conversations. Project page: https://jjjjiaozi.github.io/TS-VAP/.
CommentsSubmitted to ICASSP 2027