发表机构
University of Salerno(萨莱诺大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对音频 MLLM 可解释性,提出首个词元级时频定位框架 STAG,通过词汇投影与频谱遮挡生成相关性图,在四个基准上取得最优事件定位性能,并验证了解释的忠实性与选择性。
AI 中文摘要
基于音频的多模态大语言模型(MLLMs)能够生成复杂声学场景的详细自然语言描述,然而目前仍不清楚输入音频的哪些部分支持每个生成的词元。这一问题尤其具有挑战性,因为声学证据分布在时间和频率维度上,并且并发的声音事件可能在时间上重叠,同时占据不同的频谱区域。我们提出了 STAG,据我们所知,这是首个针对基于音频的 MLLM 生成的描述进行词元级时频定位的后置解释框架。STAG 通过使用目标词元特定的词汇投影对编码后的音频表示进行投影,来估计每个生成词元的时间支持;通过受控的频谱遮挡来测量频带相关性;并将这两个信号组合成时频相关性图。我们在四个定位基准上将 STAG 与十种后置解释方法进行了评估,它在每个数据集上都取得了最佳的事件定位性能,并将其应用于八个音频-语言骨干网络,无需参数更新。反事实删除进一步表明,移除识别出的证据会选择性地降低对相应事件的置信度,并经常将其从重新生成的描述中移除。这些结果为解释的忠实性和选择性提供了行为层面的支持。
英文摘要
Audio-based Multimodal Large Language Models (MLLMs) can generate detailed natural-language descriptions of complex acoustic scenes, yet it remains unclear which parts of the input audio support each generated token. This is particularly challenging because acoustic evidence is distributed across time and frequency, and concurrent sound events may overlap temporally while occupying different spectral regions. We introduce STAG, to our knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs. STAG estimates the temporal support for each generated token using target-token-specific vocabulary projections of the encoded audio representations, measures frequency-band relevance through controlled spectral occlusion, and combines the two signals into a spectro-temporal relevance map. We evaluate STAG against ten post-hoc explanation methods across four grounding benchmarks, where it achieves the best event-localization performance on every dataset, and apply it to eight audio-language backbones without parameter updates. Counterfactual deletion further shows that removing the identified evidence selectively reduces confidence in the corresponding event and frequently removes it from the regenerated caption. These results provide behavioral support for the faithfulness and selectivity of the explanations.