音频-视觉大型语言模型中的组合式失效:跨模态冲突下的深层先验主导
Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict
浏览论文内容
中文总结 AI 辅助
本研究针对AV-LLMs,以音频-视觉冲突为组合泛化测试,发现模型存在先验主导的组合式失效,开展可解释性分析并验证时间对齐无法提升冲突解决能力,提供了相关代码与数据。
中文摘要 AI 辅助
本研究将音频-视觉冲突作为音频-视觉大型语言模型(AV-LLMs)的组合泛化测试:模型需结合同步但语义不兼容的音频与视觉证据,判断该配对是否匹配。在VideoLLaMA 2-7B-AV模型上,三种对齐配置在AVHBench的精确字符串是/否子集中的表现接近随机水平,即便其输出先验发生了显著偏移。同样,现成的InternVideo2在跨模态冲突下准确率下降了32.3%,同时指令遵循失败率上升了17.3%。我们将这种失效模式称为先验主导:模型深层层对内部偏好的答案模式产生依赖,而该模式与冲突输入的关联较弱。为解释该行为,我们开展了机制可解释性分析,发现这种依赖集中在25.5±1层。研究还表明,更强的时间对齐会改变答案偏差,但无法提升组合冲突解决能力。可在此处的URL获取用于复现机制审计与行为评估的代码和数据。
英文摘要
We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain nearchance on the scored exact-string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off-the-shelf InternVideo2 experienced a 32.3% accuracy decrease specifically under cross-modal conflict, accompanied by a 17.3% instruction-following failure. We call this failure mode prior dominance: late-layer commitment to an internally preferred answer pattern that is weakly grounded in the conflicting inputs. To explain this behavior, we conduct a mechanistic interpretability analysis and find that commitment remains concentrated at 25.5 $\pm$ 1 layers. We show that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution. Code and data to reproduce our mechanistic audit and behavioral evaluations are available at https://github.com/AdarshSudheer09/AVHBench-dmai.