arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于最优传输的基于大语言模型的视听语音识别语义对齐

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai

arXiv 2607.09001首次发表:更新:

AI 中文总结

研究基于大语言模型的视听语音识别,提出基于最优传输的语义对齐框架,通过在多模态融合前对齐声学和视觉表征弥合模态差距,经实验验证该方法能有效提升性能,在多种条件下达到最优。

AI 中文摘要

基于大语言模型(LLM)的视听语音识别(LLM-AVSR)通过利用互补的音频和视觉信息,在不利声学环境中展现出强大的鲁棒性。现有方法通常采用独立预训练的声学和视觉编码器,其输出经投影和融合作为软提示来调节LLM进行语音识别。但多数方法未明确解决音频、视觉和文本模态间的表征差异,可能限制跨模态整合效果。本文提出基于最优传输(OT)的LLM-AVSR语义对齐框架。该方法在多模态融合前,通过参考LLM的语言嵌入空间对齐声学和视觉表征,明确弥合模态差距。具体用OT估计概率耦合矩阵,将其用作软伪标签监督对比学习,鼓励提取语义连贯和跨模态一致的视听表征。通过将多模态特征锚定到LLM的语言空间,促进更有效的多模态融合和解码。利用基于Whisper的声学编码器、基于AV-HuBERT的视觉编码器和LLaMA3.2-3B解码器实现该框架。在LRS3-TED基准上的实验表明,相较于强基线有持续改进,在各种信噪比的干净和噪声评估条件下均达到了当前最优性能。

英文摘要

Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) has recently demonstrated strong robustness in adverse acoustic environments by leveraging complementary audio and visual information. Existing approaches typically employ independently pretrained acoustic and visual encoders, whose outputs are projected and fused as soft prompts to condition an LLM for speech recognition. However, most methods perform multimodal fusion without explicitly addressing the representational discrepancy between audio, visual and text modalities, potentially limiting the effectiveness of cross-modal integration. In this paper, we propose an optimal transport (OT)-based semantic alignment framework for LLM-AVSR. The proposed method explicitly bridges the modality gap by aligning the acoustic and visual representations with reference to the linguistic embedding space of the LLM before multimodal fusion. Specifically, OT is used to estimate probabilistic coupling matrices that characterize structured correspondences between modality-specific features and linguistic embeddings. The resulting OT couplings are further utilized as soft pseudo-labels to supervise contrastive learning, encouraging the extraction of semantically coherent and cross-modal consistent audio-visual representations. By anchoring multimodal features to the linguistic space of the LLM, the proposed framework facilitates more effective multimodal fusion and decoding. We implement the proposed framework using a Whisper-based acoustic encoder, an AV-HuBERT-based visual encoder, and a LLaMA3.2-3B decoder. Experiments conducted on the LRS3-TED benchmark demonstrate consistent improvements over strong baselines and achieve state-of-the-art performance under both clean and noisy evaluation conditions across a wide range of signal-to-noise ratios (SNRs).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑