发表机构
Seoul National University; Shanghai Jiao Tong University(首尔大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AdaLoop通过轻量级循环模块自适应决定音频-问题对的推理深度,以不到3%的参数增加在多个模型和基准上将平均准确率提升2.9至3.8个百分点。
AI 中文摘要
大型音频语言模型能够回答关于语音、声音和音乐的问题,但在需要细粒度声学分析的任务上,其准确性会急剧下降。判断两个说话者中谁的音调更高,需要迭代的信号级推理,而内容性问题则不需要。当前模型在这两类任务上花费相同的计算深度。我们引入了AdaLoop,一个轻量级循环模块,它学习给定的音频-问题对需要多少潜在细化步骤。一个共享的Transformer块在问题的引导下对音频表示进行迭代,同时一个学习到的停止机制在表示准备好后退出循环。AdaLoop增加的参数不到基础模型的3%,并且可以插入任何音频编码器-语言模型对中,而无需修改任一组件。在MMSU、MMAU-Pro和MMAR上对三个架构不同的模型进行评估,AdaLoop将平均准确率提高了2.9到3.8个百分点,其中在感知密集型子任务上提升最大,模型在这些任务上学会了应用更深层次的推理。
英文摘要
Large audio language models answer questions about speech, sound, and music, yet their accuracy drops sharply on tasks that need fine-grained acoustic analysis. Judging which of two speakers has the higher pitch demands iterative signal-level reasoning that a content question does not. Current models spend the same computational depth on both. We introduce AdaLoop, a lightweight recurrent module that learns how many latent refinement steps a given audio--question pair requires. A shared transformer block iterates over the audio representation, guided by the question, while a learned halting mechanism exits the loop once the representation is ready. AdaLoop adds fewer than 3\% of the base model's parameters and plugs into any audio encoder--language model pair without modifying either component. Evaluated on three architecturally distinct models across MMSU, MMAU-Pro, and MMAR, AdaLoop raises the average accuracy by 2.9 to 3.8 points, with the largest gains on perception-heavy subtasks where the model learns to apply deeper reasoning.