arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30776eess.AS

用于基于大语言模型的自动语音识别幻觉缓解的似然约束声学重排序(无需训练)

Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR

Jiasheng Kuang, Linru Zheng, Hongjin Song, Zhaoqi Cui, Song Li

AI总结:

本研究针对基于LLM的ASR系统的幻觉问题,提出无需训练的LCAR方法,通过似然约束与声学重排序缓解幻觉,在δ=0.60时消除38.8%-57.1%的幻觉,同时维持词/字符错误率。

AI中文摘要:

基于大语言模型(LLM)的自动语音识别(ASR)系统,凭借强大的语言先验和多语言能力,在常规语音数据上取得了出色性能。然而,在具有挑战性的条件下,这些先验可能会覆盖声学证据,导致意外的翻译、指令执行、重复或灾难性删除。本文提出了似然约束声学重排序(LCAR),这是一种无需训练的解码方法,可在保留基础模型支持的同时提升声学接地能力。在每个解码步骤中,LCAR首先保留基础模型似然落在贪心令牌边际内的令牌,然后使用从注意力池化音频嵌入和现有语言模型(LM)头计算的声学兼容性得分对这些令牌进行重排序。通过将声学干预限制在合理、模型支持的备选方案中,LCAR在推理时无需额外训练、外部检测器、参考转录本或辅助模型。我们使用人工审核的TTS和开源语音挑战套件,在四个基于LLM的ASR系统上评估LCAR。当δ=0.60时,LCAR消除了检测器识别出的38.8%至57.1%的幻觉失败,同时在标准开源测试集上基本保持词错误率(WER)/字符错误率(CER)。

英文摘要:

Large language model (LLM)-based automatic speech recognition (ASR) systems achieve strong performance on conventional speech data by leveraging powerful linguistic priors and multilingual capabilities. However, under challenging conditions, these priors can override acoustic evidence, resulting in unintended translation, instruction execution, repetition, or catastrophic deletion. We propose Likelihood-Constrained Acoustic Reranking (LCAR), a training-free decoding method that improves acoustic grounding while preserving support from the base model. At each decoding step, LCAR first retains tokens whose base-model likelihood falls within a margin of the greedy token, then reranks them using an acoustic compatibility score computed from attention-pooled audio embeddings and the existing LM head. By restricting acoustic intervention to plausible, model-supported alternatives, LCAR requires no additional training, external detector, reference transcript, or auxiliary model at inference. We evaluate LCAR on four LLM-based ASR systems using human-audited TTS and open-source speech challenge suites. At $δ=0.60$, LCAR removes 38.8--57.1\% of detector-identified hallucination failures while largely maintaining WER/CER on standard open-source test sets.

补充信息

↑