发表机构
The University of Melbourne; University of Auckland; University of Birmingham; ARC OPTIMA(墨尔本大学; 奥克兰大学; 伯明翰大学; ARC OPTIMA)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出感知接地测试时强化学习(PG-TTRL),通过量化并增强大型音频语言模型推理中对声学证据的依赖,显著提升音频推理准确率,优于基础模型和标准TTRL。
AI 中文摘要
大型音频语言模型(LALMs)正越来越多地被用于更广泛的音频推理任务。这些模型通常将音频表示整合到大型语言模型(LLM)主干中,以实现多模态推理。最近的测试时强化学习(TTRL)方法通过利用预训练后的未标注测试数据,进一步提升了LLM的推理能力。然而,LALMs的感知能力的重要性仍未得到充分探索,特别是在推理过程中整合和依赖了多少声学证据,以及这对最终任务性能的贡献如何。这一空白限制了诸如TTRL等有效后训练方法在音频推理中的发展。在这项工作中,我们首先分析了音频信息在推理过程中是如何被整合和利用的。我们量化了逐层的感知依赖,并表明更强的声学依赖与更高的准确率和更大的归因于音频输入的性能提升相关联。基于此,我们提出了感知接地TTRL(PG-TTRL),该方法将无标签的测试时优化与感知接地的推理对齐,鼓励模型在推理时更强烈地基于音频输入构建推理结构。跨LALMs和基准的实验表明,PG-TTRL在推理性能上持续优于基础模型和标准TTRL,显示了感知接地优化对测试时音频推理的价值。
英文摘要
Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio representations into a large language model (LLM) backbone to enable multimodal reasoning. Recent test-time reinforcement learning (TTRL) methods further improve LLM reasoning capability by leveraging unlabelled test data after pre-training. However, the importance of the perceptual capability of LALMs remains underexplored, particularly how much acoustic evidence is integrated and relied upon during reasoning, and how this contributes to final task performance. This gap limits the development of effective post-training methods like TTRL for audio reasoning. In this work, we first analyse how audio information is integrated and utilised during reasoning process. We quantify layer-wise perceptual reliance and show that stronger acoustic reliance is associated with higher accuracy and a larger performance gain attributable to the audio input. Building on this, we propose Perception-Grounded TTRL (PG-TTRL), which aligns label-free test-time optimisation with perceptually grounded reasoning, encouraging the model to structure its reasoning more strongly on the audio input. Experiments across LALMs and benchmarks show that PG-TTRL consistently improves reasoning performance over both the base models and standard TTRL, showing the value of perceptual-grounding optimisation for test-time audio reasoning.