发表机构
Hanyang University(汉阳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对鲁棒视听语音识别问题,提出注意力引导的对比解码可靠性缩放方法,在LRS3数据集上实现了干净与低信噪比条件下的性能提升。
AI 中文摘要
基于大语言模型(LLM)的视听语音识别(AVSR)系统在噪声环境下具备鲁棒性。对比解码(CD)最初用于在推理时通过将较弱模型与较强模型进行对比来稳定LLM的生成,无需额外训练即可调整预测结果。本研究中,我们将CD应用于AVSR,在同一基础模型内将仅音频条件与完整视听条件进行对比。然而,使用固定对比强度会在不同噪声水平间引入权衡:更强的干预在严重噪声下有帮助,但在干净条件下可能过度修正可靠的预测结果。我们提出针对AVSR的CD可靠性感知缩放方法,不再使用固定强度,而是基于从注意力动态和模型间预测发散中导出的可靠性信号,自适应地调整每个token的对比影响。在LRS3数据集上的实验表明,该方法在干净和低信噪比(SNR)条件下均实现了一致的性能提升。
英文摘要
Large language model (LLM)-based audio-visual speech recognition (AVSR) systems are robust under noise. Contrastive decoding (CD), originally introduced to stabilize LLM generation by contrasting a weaker model against a stronger one at inference time, adjusts predictions without additional training. In this work, we apply CD to AVSR by contrasting audio-only conditioning with full audio-visual conditioning within the same underlying model. However, using a fixed contrastive strength introduces a trade-off across noise levels: stronger intervention helps under severe noise but may over-correct reliable predictions in clean conditions. We propose reliability-aware scaling of CD for AVSR. Instead of using a fixed strength, we adaptively modulate the contrastive influence at each token based on reliability signals derived from attention dynamics and inter-model predictive divergence. Experiments on LRS3 show consistent improvements across clean and low-SNR conditions.
CommentsAccepted to Interspeech 2026