发表机构
Worcester Polytechnic Institute(伍斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对目标说话人ASR在域偏移下性能下降的问题,提出基于Whisper的非对称分类器无关引导,通过单一引导尺度调整说话人条件贡献,并训练轻量级预测器逐句优化尺度,实现最高21.8%的相对词错误率降低。
AI 中文摘要
目标说话人自动语音识别(TS-ASR)必须在不同的重叠和噪声条件下识别并转录出期望的说话人。这些变化改变了语音混合中目标说话人的声学证据,从而激发了对说话人条件进行推理时校准的需求。我们针对基于Whisper的TS-ASR引入了非对称分类器无关引导(CFG):说话人条件分支预测目标转录,而说话人无条件分支预测序列化的多说话人转录。CFG通过单个引导尺度在解码过程中调整说话人条件的贡献。我们在目标域开发数据上选择全局引导尺度,并训练一个轻量级基于编码器的预测器来为每个话语调整该尺度,同时保持识别模型固定。在域偏移下,我们的完整系统相对于仅条件基线实现了高达21.8%的相对词错误率(WER)降低,相对于同一CFG训练模型的标准条件解码实现了5.6%的降低。Oracle分析表明,通过话语级尺度选择可以实现更大的WER降低,并识别了有益调整如何随域偏移而变化。
英文摘要
Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adjust it for each utterance, keeping the recognition model fixed. Under domain shifts, our full system achieves relative word error rate (WER) reductions of up to 21.8% over the condition-only baseline, and 5.6% over standard conditional decoding of the same CFG-trained model. Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.