多模态目标说话人提取:迈向跨模态统一说话人线索
Multimodal Target Speaker Extraction: Towards Unified Speaker Cues Across Modalities
- University of Science and Technology Beijing(北京科技大学)
- Beijing Institute of Technology(北京理工大学)
- Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Alibaba Group(阿里巴巴集团)
- Zhejiang Institute of Quality Sciences(浙江省质量科学研究院)
- Tianjin University(天津大学)
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文综述多模态目标说话人提取,分类五类线索,梳理模型演变,讨论挑战并展望融合、指令驱动等未来方向。
AI中文摘要:
目标说话人提取(TSE)在语音通信和人机交互中至关重要,它能够从复杂的声学环境(即鸡尾酒会场景)中隔离出特定说话人的声音。尽管基于注册语音的传统TSE系统已取得显著进展,但注册语音作为线索存在固有局限性。当目标说话人与干扰说话人具有相似嗓音特征时,当说话人内部变异性(如情绪或说话风格的变化)导致注册语音与目标语音不匹配时,或当注册语音本身被噪声或竞争说话人污染时,其可靠性会下降。本综述从多模态辅助目标线索的角度,对基于深度学习的TSE进行了调研。我们根据用于隔离目标说话人的五种信息类型对现有方法进行归类:音频注册、视觉、空间、文本/语义和神经线索。我们还追溯了从判别式估计器到基于变分、扩散、流、编解码器和基础模型系统的演变,并总结了代表性数据集和评估指标。我们回顾了不同线索的优势与局限,并讨论了涉及同步、缺失或不可靠观测、数据稀缺、隐私、计算成本和实时操作方面的挑战。最后,我们总结了关于自适应线索融合、指令驱动提取、现实评估和可信部署的未来方向。通过联合审视线索设计、模型架构、训练目标、数据集和评估指标,本文概述了多模态TSE的当前格局和开放问题。
英文摘要:
Target Speaker Extraction (TSE) is pivotal in speech communication and human-computer interaction, enabling the isolation of a specific speaker's voice from complex acoustic environments, i.e., the cocktail party scenario. Although traditional TSE systems conditioned on enrollment speech have progressed substantially, enrollment speech as a cue has inherent limitations. Its reliability degrades when the target and interfering speakers have similar voice characteristics, when intra-speaker variability (e.g. changes in emotion or speaking style) creates a mismatch between the enrollment and target speech, or when the enrollment itself is contaminated by noise or competing speakers. This review surveys deep-learning-based TSE from the perspective of auxiliary target cues drawn from multiple modalities. We organize existing methods according to five types of information used to isolate the target speaker: audio enrollment, visual, spatial, textual/semantic, and neural cues. We also trace the evolution from discriminative estimators to variational, diffusion, flow, codec, and foundation-model-based systems and summarize representative datasets and evaluation metrics. We review the benefits and limitations of different cues and discuss challenges involving synchronization, missing or unreliable observations, data scarcity, privacy, computational cost, and real-time operation. Finally, we summarize future directions concerning adaptive cue fusion, instruction-driven extraction, realistic evaluation, and trustworthy deployment. By jointly reviewing cue design, model architecture, training objectives, datasets, and evaluation metrics, this article provides an overview of the current landscape and open problems in multimodal TSE.