arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25355cs.SD

从语义到读出:时间音频接地微调后音频令牌的机制理解

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding

Yujian Ma, Jinqiu Sang, Ruizhe Li, Jiaao Yu, Ang Li

首次发表
浏览论文内容

中文总结 AI 辅助

研究大型音频语言模型中音频令牌在微调后的机制,通过四种分析研究其分层语义等,发现微调前已有潜在证据,微调后解码器更易访问事件信息,支持从语义到读出的解释,有助于解码器连接时间输出。

中文摘要 AI 辅助

大型音频语言模型(LALMs)通过原生音频令牌向语言解码器传达声学证据,但其内部作用尚不清楚。本文以时间音频接地为诊断设置,通过四种互补分析,研究语言模型微调如何影响原生音频令牌状态的分层语义、解码器可及性和时间输出对齐。实验表明,微调前查询事件的潜在证据已存在,微调后音频令牌中与查询事件最强烈对齐的部分出现在相似时间位置,事件相关信息对解码器更易访问,主要源于解码器适应。时间探针显示基础检查点已包含可恢复信息,微调主要改善与各检查点自身预测时间支持的对齐。残差增量擦除表明,在预测窗口内移除音频令牌更新对时间戳生成的损害大于移除相同数量随机选择的更新。这些结果支持从语义到读出的解释,即接地微调有助于解码器读取现有事件证据并更可靠地连接到时间输出。

英文摘要

Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Using temporal audio grounding as a diagnostic setting, we examine how language-model fine-tuning affects the layerwise semantics, decoder accessibility, and temporal output alignment of native audio-token states through four complementary analyses: query-conditioned token semantics, calibrated token readout, temporal-window probes, and residual-delta erasure during generation. Alongside substantial improvements in temporal localization, semantic analysis of Qwen2.5-Omni shows that latent evidence for queried events is already present before fine-tuning and that the audio tokens most strongly aligned with the queried event appear at similar temporal positions before and after fine-tuning. After fine-tuning, event-related information in audio tokens becomes more accessible to the decoder, especially in early and middle layers, and a cross-checkpoint control shows that this improvement arises primarily from decoder adaptation. Temporal probes show that the base checkpoint already contains recoverable information about annotated windows and that fine-tuning mainly improves alignment with each checkpoint's own predicted temporal support. Residual-delta erasure further shows that removing audio-token updates within predicted windows harms timestamp generation more than removing the same number of randomly selected updates. The same broad improvements in decoder readability and prediction alignment also appear in Qwen2-Audio. Together, these results support a semantics-to-readout account in which grounding fine-tuning helps the decoder read existing event evidence and connect it more reliably to temporal outputs.

↑