以声音为锚:利用冻结声学评判器进行强化学习以抑制ASR插入幻觉
Grounded in Sound: Reinforcement Learning with a Frozen Acoustic Judge to Curb ASR Insertion Hallucinations
查看机构详情
- Tencent(腾讯)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对ASR后训练中RL奖励仅基于文本导致插入幻觉的问题,提出用冻结的wav2vec2-CTC声学评判器增强GRPO奖励,在AMI上减少插入错误28.3%并降低WER,且推理时无需额外开销。
中文摘要 AI 辅助
当强化学习(RL)用于自动语音识别(ASR)的后训练时,奖励几乎总是存在于文本空间中:它将假设与参考进行比较,从不检查假设是否得到音频的支持。在高度规则的语音上,这允许一种捷径——从强大的语言先验中猜测,而不是倾听。一旦声学条件退化,捷径便不受约束地运行,产生流畅但无根据的词,即插入错误。我们提出一种声学保真度奖励:一种GRPO奖励,附加一个单独预训练、永久冻结、非自回归的字符级wav2vec2-CTC声学评判器,仅在训练时使用,推理时不存在,推理时单个模型进行贪心解码。在LibriSpeech上训练,并在包括真实AMI会议语音(33,282个话语条件实例)的六级难度梯度上评估,该方法在近距离拾音AMI-IHM上将插入错误减少28.3%,在远场AMI-SDM上减少22.3%,同时将AMI-SDM上的WER从35.89%降至34.71%,与其他五级相比,WER无显著差异,对照为调度匹配的WER-GRPO基线。插入减少在会议级聚类自助法下保持稳健。四项预设分析支持内容条件插入校准:在保留能量和语音活动的不可理解音频上,输出坍缩85-90%;评估的32最佳CTC重打分配置无法恢复该增益,但RL将其内化到单次贪心解码运行中;仅策略置信度在所有四个评估设置中产生较低的插入AURC。我们将其定位为一篇机制论文,在一个实例中演示:一个7B语音大语言模型配0.3B CTC评判器。
英文摘要
When reinforcement learning (RL) is used for post-training automatic speech recognition (ASR), the reward almost always lives in the text space: it compares a hypothesis with the reference and never checks whether the hypothesis is supported by the audio. On highly regular speech this licenses a shortcut - guessing from a strong language prior rather than listening. Once the acoustics degrade, the shortcut runs unchecked and emits fluent but ungrounded words, i.e., insertion errors. We propose an acoustic-fidelity reward: a GRPO reward augmented with a separately pretrained, permanently frozen, non-autoregressive character-level wav2vec2-CTC acoustic judge, used strictly at training and absent at inference, where a single model decodes greedily. Trained on LibriSpeech and evaluated across a six-tier difficulty gradient including real AMI meeting speech (33,282 utterance-condition instances), the method reduces insertion errors by 28.3% on close-talking AMI-IHM and 22.3% on far-field AMI-SDM, while lowering WER on AMI-SDM from 35.89% to 34.71% and showing no detectable WER difference on the other five tiers, against a schedule-matched WER-GRPO baseline. The insertion reduction holds under a meeting-level clustered bootstrap. Four prespecified analyses support content-conditioned insertion calibration: output collapses 85-90% on unintelligible audio that preserves energy and voice activity; the gain is not recovered by the evaluated 32-best CTC rescoring configuration, yet RL internalizes it into a single greedy decoding run; and policy-only confidence yields lower insertion-AURC in all four evaluated settings. We frame this as a mechanism paper, demonstrated in one instantiation: a 7B speech LLM with a 0.3B CTC judge.