发表机构
Korea Advanced Institute of Science and Technology(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探究Audio LLMs中音频决定答案时的内部变化,通过实验发现训练后模型对音频替换更敏感,声学信息在不同层的作用及训练权重的影响,为其使用声学证据提供机制解释。
AI 中文摘要
音频大语言模型(Audio LLMs)在音频理解方面已取得进展,但仍可能通过文本线索或语言先验而非提供的音频来预测答案。常见的解决方法是在答案无法仅从文本推断的数据上训练模型,该方法可提升性能,但模型内部的变化仍不明确。本文探究音频实际决定答案时模型内部必须发生的变化,研究发现有三点:(1)将音频替换为静音或不相关音频时,训练后的模型性能下降幅度远大于预训练模型;(2)声学信息在早至中层对模型答案选项表征的塑造作用最强,而训练主要在中至后期层增强音频信息对最终预测的影响;(3)训练期间学习到的权重在特定层段影响最大。综上,这些结果为训练如何加强Audio LLMs中声学证据的使用提供了机制性解释。
英文摘要
Audio Large Language Models (Audio LLMs) have advanced in audio understanding, yet they can still predict the answer by reasoning from textual cues or linguistic priors rather than the provided audio. A common remedy is to train models on data whose answers cannot be inferred from text alone. This approach can improve performance, but what changes within the model remains unclear. In this paper, we ask what must happen inside the model for the audio to actually determine the answer. Our findings are threefold. (1) Replacing the audio with silence or unrelated audio causes substantially larger performance degradation in the trained model than in the pretrained model. (2) Acoustic information most strongly shapes the model's representations of the answer choices in early-to-middle layers, while training mainly increases the influence of audio information on the final prediction in middle-to-late layers. (3) The weights learned during training have their largest impact in specific layer bands. Together, these results provide a mechanistic account of how training strengthens the use of acoustic evidence in Audio LLMs.
CommentsPreprint