统一目标说话人语音识别:基于文本与注册语音线索
Unified Target-Speaker ASR with Text and Enrollment Speech Cues
浏览论文内容
中文总结 AI 辅助
提出统一双线索TS-ASR框架,融合文本与注册语音线索,在30,000混合语音上实现8.80% CER,优于单一线索方法。
中文摘要 AI 辅助
目标说话人自动语音识别(TS-ASR)旨在多说话人环境中识别指定说话人,同时抑制干扰语音。传统TS-ASR通常依赖注册语音,而文本引导方法利用已知词汇内容(如唤醒词)从观测混合语音中识别目标说话人。这两种线索提供互补信息,但通常被分开研究。我们提出一种统一双线索TS-ASR框架,在单一模型中支持文本线索、注册语音或两者兼用。文本线索与混合语音表示交互,根据已知词汇内容提取目标说话人信息,而独立的注册语音提供互补的说话人信息。跨注意力线索条件模块集成到共享的Conformer块中,负线索采样在双线索训练期间提供线索有效性监督。在包含五种录音/领域条件和四种真实文本线索长度的30,000个双说话人混合语音上的实验表明,使用五字符文本线索时,拼接式双线索方法达到8.80%的CER,而仅文本和仅注册语音推理分别为17.32%和29.06%。该方法还优于并行双线索融合(9.49% CER),并在所有五个评估子集上获得更低的双线索CER。这些结果证明了联合利用互补的词汇和说话人信息对目标说话人ASR的益处。
英文摘要
Target-speaker automatic speech recognition (TS-ASR) aims to recognize a designated speaker while suppressing interfering speech in multi-talker environments. Conventional TS-ASR typically relies on an enrollment utterance, whereas text-guided methods use known lexical content, such as a wake word, to identify the target speaker from the observed mixture. These two cues provide complementary information but are usually studied separately. We propose a Unified Dual-Cue TS-ASR framework that supports text cues, enrollment speech, or both within a single model. Text cues interact with the mixture representation to extract target-speaker information conditioned on known lexical content, while an independent enrollment utterance provides complementary speaker information. Cross-attention cue-conditioning modules are integrated into shared Conformer blocks, and negative-cue sampling provides cue-validity supervision during dual-cue training. Experiments on 30,000 two-speaker mixtures across five recording/domain conditions and four oracle text-cue lengths show that, with five-character text cues, the concatenated dual-cue method achieves 8.80% CER, compared with 17.32% for text-only and 29.06% for enrollment-only inference. It also outperforms parallel dual-cue fusion (9.49% CER) and yields lower dual-cue CER across all five evaluation subsets. These results demonstrate the benefit of jointly exploiting complementary lexical and speaker information for target-speaker ASR.
发表机构
- Shanghai Normal University(上海师范大学)
- Baidu(百度)
- Unisound(云知声)
机构由 AI 辅助整理,请以论文原文为准。