arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaLoop:音频语言模型的自适应深度潜在推理

AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models

Lee Seung-woo, Bowen Qi

arXiv 2610.06949首次发表:更新:

发表机构

Seoul National University; Shanghai Jiao Tong University(首尔大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AdaLoop通过轻量级循环模块自适应决定音频-问题对的推理深度,以不到3%的参数增加在多个模型和基准上将平均准确率提升2.9至3.8个百分点。

AI 中文摘要

大型音频语言模型能够回答关于语音、声音和音乐的问题,但在需要细粒度声学分析的任务上,其准确性会急剧下降。判断两个说话者中谁的音调更高,需要迭代的信号级推理,而内容性问题则不需要。当前模型在这两类任务上花费相同的计算深度。我们引入了AdaLoop,一个轻量级循环模块,它学习给定的音频-问题对需要多少潜在细化步骤。一个共享的Transformer块在问题的引导下对音频表示进行迭代,同时一个学习到的停止机制在表示准备好后退出循环。AdaLoop增加的参数不到基础模型的3%,并且可以插入任何音频编码器-语言模型对中,而无需修改任一组件。在MMSU、MMAU-Pro和MMAR上对三个架构不同的模型进行评估,AdaLoop将平均准确率提高了2.9到3.8个百分点,其中在感知密集型子任务上提升最大,模型在这些任务上学会了应用更深层次的推理。

英文摘要

Large audio language models answer questions about speech, sound, and music, yet their accuracy drops sharply on tasks that need fine-grained acoustic analysis. Judging which of two speakers has the higher pitch demands iterative signal-level reasoning that a content question does not. Current models spend the same computational depth on both. We introduce AdaLoop, a lightweight recurrent module that learns how many latent refinement steps a given audio--question pair requires. A shared transformer block iterates over the audio representation, guided by the question, while a learned halting mechanism exits the loop once the representation is ready. AdaLoop adds fewer than 3\% of the base model's parameters and plugs into any audio encoder--language model pair without modifying either component. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, AdaLoop raises the average accuracy by 2.9 to 3.8 points, with the largest gains on perception-heavy subtasks where the model learns to apply deeper reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑