发表机构
Portland State University; US Army Research Laboratory(波特兰州立大学; 美国陆军研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对级联多说话人自动语音识别中的说话人泄漏问题,提出基于剪枝的校正范式,利用预训练说话人 diarization 模型剪枝转录片段,在多语料库上实现 cpW ER 最高相对降低 29%,提升了转录可靠性。
AI 中文摘要
级联多说话人自动语音识别(MT-ASR)虽利用了最先进的基础模型,但其性能常受分离过程中的说话人泄漏限制。现有校正策略主要聚焦于针对说话人归属的词汇重标注,我们提出一种互补的基于剪枝的范式,可鲁棒识别并去除泄漏伪影。该方法利用预训练的说话人 diarization 模型作为多模态验证器,对满足时间包含性、词汇交叉验证及时间对齐三方共识的转录片段进行剪枝。在 LibriMix、LibriSpeechMix 和 AMI Meeting 语料库上的结果显示,我们的算法在不同重叠条件下均能持续降低 cpW ER,具体而言,在高说话人泄漏的子集上,该方法实现了 cpW ER 相对降低最高达 29%,凸显了其在复杂声学环境中提升级联 MT-ASR 转录可靠性的有效性。
英文摘要
While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during separation. Prior correction strategies primarily focus on lexical re-labeling for speaker attribution. We propose a complementary pruning-based paradigm that robustly identifies and removes leakage artifacts. Our method utilizes a pre-trained speaker diarization model as a multimodal verifier to prune transcribed segments satisfying a tripartite consensus of temporal containment, lexical cross-validation, and temporal alignment. Results on LibriMix, LibriSpeechMix, and the AMI Meeting corpus show our algorithm consistently reduces cpW ER across diverse overlap conditions. Specifically, on subsets with high speaker leakage, our method achieves relative cpW ER reductions of up to 29%, highlighting its effectiveness in enhancing the reliability of cascaded MT-ASR transcripts in complex acoustic environments.
CommentsAccepted to INTERSPEECH 2026