arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27026cs.SDcs.MM

直接还是中介?大型音频语言模型中任务依赖的音频信息路由

Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models

  • Graduate School of Informatics, Kyoto University(京都大学信息学研究科)
  • WXG, Tencent(腾讯微信事业群)

机构由 AI 辅助整理,请以论文原文为准。

Yizhou Zhang, Wangjin Zhou, Xin Gu, Yichi Wang, Wei Tan, Yi Zhao, Zhi Gong, Keisuke Imoto, Tatsuya Kawahara

AI总结:

该研究针对大型音频语言模型,发现其在拼接两段音频的任务中,自动语音识别稳定而音频问答性能大幅下降,揭示了两类任务依赖不同的音频信息路由路径,指出信息利用是其泛化的潜在限制。

AI中文摘要:

大型音频语言模型(Large Audio Language Models, LALMs)在各类音频理解任务中展现出强劲性能,但它们通常仅在单一连贯音频片段上进行评估,对其在不熟悉输入配置下的表现研究不足。我们通过将两段音频片段拼接为单个输入的受控设置研究该问题,在多个LALMs中观察到显著的任务依赖鲁棒性差距:自动语音识别(Automatic Speech Recognition, ASR)表现相对稳定,而音频问答(Audio Question Answering, AQA)性能大幅下降。为探究该差异的潜在机制,我们采用分层注意力敲除法分析音频信息如何通过LALMs解码器路由,结果揭示了不同的任务依赖路径:ASR主要依赖答案令牌从音频令牌直接检索信息,而AQA更依赖中介路径,即音频信息先整合到提示令牌中,再在生成过程中被访问。我们进一步探查音频拼接下的提示令牌表示,发现即使AQA性能急剧下降,任务相关的音频属性仍可被轻松解码,尤其在解码器的中间层和后期层。这种分离表明,该故障无法用解码器状态中音频信息的完全丢失来解释,而是与答案生成过程中检索或利用提示中介信息的下游瓶颈一致。综上,我们的发现揭示了LALMs中任务依赖的音频信息路由,并强调信息利用是其泛化能力的潜在限制因素。

英文摘要:

Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio understanding tasks. However, they are typically evaluated on single, coherent audio segments, leaving their behavior under less familiar input configurations underexplored. We study this issue through a controlled setting in which two audio segments are concatenated into a single input. Across multiple LALMs, we observe a striking task-dependent robustness gap: automatic speech recognition (ASR) remains comparatively stable, whereas audio question answering (AQA) degrades substantially. To investigate the mechanisms underlying this disparity, we analyze how audio information is routed through LALM decoders using layer-wise attention knockout. The results reveal distinct task-dependent pathways. ASR relies primarily on direct retrieval from audio tokens by answer tokens, whereas AQA depends more strongly on a mediated route in which audio information is first integrated into prompt tokens and subsequently accessed during generation. We further probe prompt-token representations under audio concatenation and find that task-relevant audio attributes remain readily decodable, particularly in middle and later decoder layers, even when AQA performance deteriorates sharply. This dissociation indicates that the failure cannot be explained by complete loss of audio information from the decoder states and is instead consistent with a downstream bottleneck in retrieving or utilizing prompt-mediated information during answer generation. Together, our findings reveal task-dependent audio information routing in LALMs and highlight information utilization as a potential limitation on their generalization.

补充信息

↑