arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ASR幻觉的剖析

The Anatomy of an ASR Hallucination

Hamees Sayed, Apoorv Singh, Kumar Aman, Akshat Mandloi

arXiv 2609.04404首次发表:更新:

发表机构

Smallest AI; Fast Code AI(Smallest AI; Fast Code AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究分析ASR幻觉的机制前提,通过对比CTC和RNN-T的Conformer-Large模型,发现最终编码器阶段是接地识别的关键边界,其失效会引发输出分歧。

AI 中文摘要

ASR系统有时会生成与接收语音无关的流畅文本,我们将这些幻觉视为更广泛的接地失效的一种可能后果,其中转录文本不再受到音频的充分引导。为理解这种失效何时可能发生,我们研究了两个独立训练的Conformer-Large识别器——一个基于CTC,另一个基于RNN-T——在环境退化和说话人背景偏移下的表现。在这两个模型中,最终编码器阶段成为关键边界:绕过最终块会导致几乎所有话语出现分歧,而绕过中间块几乎没有影响。在同一阶段,表征变得更紧凑,文本可被训练过的解码器读取,字素信息变得明确。重要的是,这种干预会产生混乱或重复的输出,而非流畅的生成。因此,我们的结果确定了幻觉的一个机制前提——生成充分接地输出的失效——而非自然发生幻觉的完整起源。这些结果共同揭示了在两种解码器家族和多种分布偏移下,接地识别存在一致的终端阶段依赖性。

英文摘要

ASR systems sometimes produce fluent text that is unrelated to the speech they receive. We view these hallucinations as one possible consequence of a broader grounding failure, in which the transcript is no longer adequately guided by the audio. To understand where this failure becomes possible, we study two independently trained Conformer-Large recognizers - one CTC and one RNN-T - under environmental degradation and speaker-background shift. In both models, the final encoder stage emerges as a critical boundary: bypassing the final block causes divergence on nearly every utterance, whereas bypassing middle blocks has little effect. At this same stage, the representations become more compact, text becomes readable by the trained decoder, and grapheme information becomes explicit. Importantly, the intervention produces garbled or repetitive output rather than fluent fabrication. Our result therefore identifies a mechanistic precondition for hallucination - the failure to produce adequately grounded output - not the complete origin of naturally occurring hallucinations. Together, the results reveal a consistent terminal-stage dependency for grounded recognition across two decoder families and multiple distribution shifts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑