发表机构
Zhejiang Lab; Zhejiang International Studies University(浙江实验室; 浙江外国语大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出软后验说话人注入(SPSI)方法,通过将帧级说话人后验注入 Whisper,在 LibriSpeech、LibriCSS 多说话人语音识别任务上降低了 cpWER,表现优于序列化输出训练(SOT)。
AI 中文摘要
多说话人自动语音识别(MT-ASR)在语音重叠场景下仍具挑战性。基于硬说话人 diarization 的分割会引入不可逆误差,而序列化输出训练(SOT)虽避免了显式分割,但未让预训练编码器适配说话人活动。我们提出软后验说话人注入(SPSI):一个轻量头部预测帧级说话人后验 $\boldsymbol{\tilde{P}}$,并通过多层特征级线性调制(FiLM)和解码器说话人记忆提示将其注入 Whisper 模型。在受控双说话人 LibriSpeech 重叠场景中,SPSI 将受话语均值约束的置换词错误率(cpWER)从 SOT 的 50.7% 降至 49.6%(单侧配对自助法 $p{\text{≈}}0.006$),高重叠区间的降幅更大(从 60.4% 降至 58.8%)。相同主干的说话人辅助目标和语音活动检测(VAD)管道未优于 SOT;零样本(ZS)LibriCSS 表现相当。采用重重叠(OV-heavy)数据继续的冻结后验适配,将保留的 LibriCSS(会话 8-9)cpWER 降至 32.4%(对比 SOT 的 37.5%)。消融实验表明,编码器 FiLM 与解码器提示具有互补作用,且有效信号是软单纯形值的说话人份额。
英文摘要
Multi-talker automatic speech recognition (MT-ASR) remains challenging in the presence of overlapping speech. Hard segmentation introduces irreversible errors, whereas serialized output training (SOT) avoids explicit segmentation but does not condition a pretrained encoder on speaker activity. We propose Soft Posterior Speaker Injection (SPSI). A Soft Posterior Head predicts per-frame speaker posteriors $\hat{\mathbf{P}}$ and injects them into Whisper through Multi-layer Feature-wise Linear Modulation (MFLM) and Speaker Memory Prompts (SMP). The benefit of SPSI is largest where overlap is heaviest and under domain transfer. On controlled two-speaker LibriSpeech overlap, SPSI reduces concatenated minimum-permutation word error rate (cpWER) from $61.5\%$ to $60.0\%$ in the high-overlap bin, and from $51.9\%$ to $51.0\%$ on the full set, relative to SOT. By contrast, Speaker CE, SD-CTC, SA-DiCoW, and Pipeline (oracle/est.\ VAD) do not outperform SOT. Freeze-posterior overlap-heavy adaptation reduces held-out LibriCSS cpWER from $42.3\%$ to $36.8\%$ on sessions $8$--$9$, a $5.5$-point gain over SOT. The source code is available at https://github.com/HackerHyper/SPSI.
CommentsThis paper is submitted to ICASSP2027