arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越文本思维链:面向大型音频语言模型的JEPA条件潜在推理

Beyond Textual Chain-of-Thought: JEPA-Conditioned Latent Reasoning for Large Audio Language Models

Donghang Wu, Haoyang Zhang, Yizhou Peng, Shreyas Gopal, Yi-Wen Chao, Chen Chen, Hexin Liu, William Tjhi, Eng-Siong Chng

arXiv 2609.34407首次发表:更新:

发表机构

Nanyang Technological University; AI Singapore; Peking University(南洋理工大学; 新加坡人工智能研究院; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出JELAR,一种基于JEPA的潜在推理框架,利用WavJEPA的声学表示指导大型音频语言模型的潜在推理,在MMAU-mini和MMAR上分别提升2.70和9.10个百分点,有效替代显式文本思维链。

AI 中文摘要

显式文本思维链(CoT)提升了大型音频语言模型(LALMs)的推理能力。然而,文本思维链通常基于音频的文本描述构建,对声学证据的访问有限,这可能会引发幻觉等问题。为解决这一模态差距问题,我们提出了JELAR,一种基于联合嵌入预测架构(JEPA)的潜在推理框架,该框架将潜在推理监督条件化于从原始波形中学习到的声学表示上。在训练期间,冻结的WavJEPA模型提供直接从原始波形学习到的表示。一个非因果专家首先构建答案感知查询,这些查询交叉关注WavJEPA嵌入以生成潜在推理目标。LALM被训练在生成其响应之前预测这些目标。实验结果表明,JELAR在MMAU-mini和MMAR上分别将Audio-Reasoner基线提升了2.70和9.10个绝对百分点,证明了JEPA条件潜在推理作为显式文本思维链监督的替代方案的有效性。

英文摘要

Explicit textual Chain-of-Thought (CoT) has improved the reasoning ability of large audio language models (LALMs). However, textual CoTs are often constructed from text captions of audio and provide limited access to the acoustic evidence, which can introduce problems like hallucination. To address this modality-gap issue, we introduce JELAR, a Joint-Embedding Predictive Architecture (JEPA)-based latent reasoning framework that conditions latent reasoning supervision on acoustic representations learned from raw waveforms. During training, a frozen WavJEPA model provides representations learned directly from raw waveforms. A non-causal expert first constructs answer-aware queries, which cross-attend to WavJEPA embeddings to produce latent reasoning targets. The LALM is trained to predict these targets before generating its response. Experimental results show that JELAR improves the Audio-Reasoner baseline by 2.70 and 9.10 absolute percentage points on MMAU-mini and MMAR, respectively, demonstrating the effectiveness of JEPA-conditioned latent reasoning as an alternative to explicit textual CoT supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑