arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于鲁棒视听语音识别中对比解码的注意力引导可靠性缩放

Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition

YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang

arXiv 2608.26213首次发表:更新:

发表机构

Hanyang University(汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对鲁棒视听语音识别问题,提出注意力引导的对比解码可靠性缩放方法,在LRS3数据集上实现了干净与低信噪比条件下的性能提升。

AI 中文摘要

基于大语言模型(LLM)的视听语音识别(AVSR)系统在噪声环境下具备鲁棒性。对比解码(CD)最初用于在推理时通过将较弱模型与较强模型进行对比来稳定LLM的生成,无需额外训练即可调整预测结果。本研究中,我们将CD应用于AVSR,在同一基础模型内将仅音频条件与完整视听条件进行对比。然而,使用固定对比强度会在不同噪声水平间引入权衡:更强的干预在严重噪声下有帮助,但在干净条件下可能过度修正可靠的预测结果。我们提出针对AVSR的CD可靠性感知缩放方法,不再使用固定强度,而是基于从注意力动态和模型间预测发散中导出的可靠性信号,自适应地调整每个token的对比影响。在LRS3数据集上的实验表明,该方法在干净和低信噪比(SNR)条件下均实现了一致的性能提升。

英文摘要

Large language model (LLM)-based audio-visual speech recognition (AVSR) systems are robust under noise. Contrastive decoding (CD), originally introduced to stabilize LLM generation by contrasting a weaker model against a stronger one at inference time, adjusts predictions without additional training. In this work, we apply CD to AVSR by contrasting audio-only conditioning with full audio-visual conditioning within the same underlying model. However, using a fixed contrastive strength introduces a trade-off across noise levels: stronger intervention helps under severe noise but may over-correct reliable predictions in clean conditions. We propose reliability-aware scaling of CD for AVSR. Instead of using a fixed strength, we adaptively modulate the contrastive influence at each token based on reliability signals derived from attention dynamics and inter-model predictive divergence. Experiments on LRS3 show consistent improvements across clean and low-SNR conditions.

CommentsAccepted to Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑