增强音频推理:通过语义摘要预测
Enhancing Audio Reasoning via Semantic Summary Prediction
- Concordia University(康考迪亚大学)
- Mila - Quebec AI Institute(米拉-魁北克人工智能研究所)
- Université Laval(拉瓦尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LALMs推理时注意力偏离音频的问题,提出SPARE方法,通过寄存器令牌与语义目标对齐,在不增加推理成本下提升零样本音频推理性能。
AI中文摘要:
大型音频语言模型(LALMs)在复杂问答任务上表现良好,但常表现出推理差距,即显式思维链(CoT)相比直接回答会降低准确率。我们假设长推理序列会将注意力从音频输入上转移开。为解决此问题,我们提出SPARE(音频推理的语义预测),该方法引入一个寄存器令牌,通过余弦相似度损失与Sentence-BERT嵌入对齐最终结论。这会在推理开始前,用目标语义目标调节模型的潜在空间。在MMAU和MMAR上使用SALMONN进行的实验表明,零样本推理能力得到提升,且对音频的早期注意力更强,且不增加额外推理成本。
英文摘要:
Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) reduces accuracy compared to direct answers. We hypothesize that long reasoning sequences shift attention away from the audio input. To address this, we propose SPARE (Semantic Prediction for Audio REasoning), which introduces a register token aligned with the final conclusion using a cosine similarity loss with a Sentence-BERT embedding. This conditions the model's latent space with the target semantic goal before reasoning begins. Experiments on MMAU and MMAR with SALMONN show improved zero-shot reasoning and stronger early attention to audio without additional inference cost.