论大型音频语言模型中的时间绑定
On Temporal Binding in Large Audio Language Models
浏览论文内容
中文总结 AI 辅助
本研究通过机制可解释性分析三个开源大型音频语言模型,发现时间信息在中间层绑定到事件名称表示,并沿低维轨迹编码,操纵该轨迹可改变前后判断,但精确时间定位依赖不同机制。
中文摘要 AI 辅助
对音频录音的时间结构进行推理,要求大型音频语言模型(LALMs)将声音事件与其时间位置相关联。理解其底层机制是诊断故障和识别可能需要改进的模型组件的第一步。利用机制可解释性,我们研究了三个开源LALMs中时间信息如何被表示并绑定到声音事件上。我们发现,在所有三个模型中,事件特定的位置信息在中间模态集成层中集中于事件名称表示中。这些表示沿着一条低维、弯曲的相对时间轨迹编码粗略的事件位置。沿此轨迹操纵事件名称表示会系统地改变“之前/之后”的判断,为这些表示有助于粗略时间推理提供了证据。相比之下,相同的干预措施并不能可靠地改变预测的起始时间戳,这表明粗略时间推理和精确的度量事件定位依赖于不同的机制。
英文摘要
Reasoning about temporal structure of audio recordings requires Large Audio Language Models (LALMs) to associate sound events with their temporal position. Understanding the underlying mechanisms is a first step toward diagnosing failures and identifying model components that may need improvement. Using mechanistic interpretability, we investigate how temporal information is represented and bound to sound events in three open-source LALMs. We find that across all three, event-specific location becomes concentrated in event name representations at intermediate modality integration layers. These representations encode coarse event position along a low-dimensional, curved relative time trajectory. Steering event name representations along this trajectory systematically shifts before/after beliefs, providing evidence that these representations contribute to coarse temporal reasoning. In contrast, the same interventions do not reliably shift predicted onset timestamps, suggesting that coarse temporal reasoning and precise metric event localization rely on distinct mechanisms.