arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04561cs.AI

通过幻觉空间投影减少Whisper中的幻觉转录文本

Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection

  • The university of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)

机构由 AI 辅助整理,请以论文原文为准。

Maryam Abbasihafshejani, Murtuza Jadliwala

AI总结:

该研究提出一种推理阶段的低秩激活投影方法,通过估计非语音校准数据的幻觉相关子空间,两种变体可大幅降低Whisper的幻觉转录率,且在抑制幻觉与语音识别性能间实现可控权衡。

AI中文摘要:

Whisper是一款广泛应用的自动语音识别(ASR)基础模型,但其生成式解码器会对几乎不含语音或完全不含语音的输入生成流畅的幻觉转录文本。我们提出一种无需训练、仅在推理阶段使用的方法,通过对解码器激活值进行低秩投影来减少这类幻觉。该方法从非语音校准数据中估计出一个与幻觉相关的紧凑子空间,推理阶段将解码器隐藏状态向该子空间的正交补空间投影。我们评估了两种变体:始终激活式,即对所有输入应用投影;门控式,即仅当Whisper预测输入可能为非语音时才应用投影。在非语音基准测试中,始终激活式投影将平均幻觉率(HR)从31.31%降至2.44%,相对降低92.21%;门控式投影将HR降至3.74%,相对降低88.05%,同时降低了对真实语音的错误拒绝率。在LibriSpeech数据集上,门控式投影使绝对词错误率(WER)上升0.33至4.39个百分点,在不同模型和拆分设置下的错误拒绝率(FRR)为0.41%至9.97%。这些结果表明,低秩激活投影可在不重新训练的情况下大幅抑制Whisper的幻觉,同时在幻觉抑制与语音识别性能之间提供可控的权衡。

英文摘要:

Whisper is a widely used foundation model for automatic speech recognition (ASR), but its generative decoder can produce fluent hallucinated transcripts for inputs containing little or no speech. We propose a training-free, inference-time method to reduce these hallucinations using low-rank projection of decoder activations. A compact hallucination-associated subspace is estimated from non-speech calibration data, and decoder hidden states are projected away from this subspace during inference. We evaluate two variants: always-on, which applies projection to all inputs, and gated, which applies it only when Whisper predicts that an input is likely non-speech. Across non-speech benchmarks, always-on projection reduces average hallucination rate (HR) from 31.31% to 2.44%, a 92.21% relative reduction, while gated projection reduces HR to 3.74%, an 88.05% relative reduction, with lower false rejection of genuine speech. On LibriSpeech, gated projection increases absolute word error rate (WER) by 0.33-4.39 percentage points and yields false-rejection rates (FRR) of 0.41--9.97% across model and split settings. These results show that low-rank activation projection can substantially suppress Whisper hallucinations without retraining, while providing a controllable trade-off between hallucination suppression and speech recognition performance.

↑