arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码但断开:在修补空值下分解视觉语言模型的失败

Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null

Genpei Zhang

arXiv 2610.00024首次发表:更新:

发表机构

University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在三种视觉语言模型中发现中层编码答案但因果断开,通过修补实验验证失败模式,并证明其可学习性与干预效果。

AI 中文摘要

在三种视觉语言模型架构(LLaVA-1.5-7B、Qwen2.5-VL-7B、InternVL3-8B)中,我们报告了一个关于中层可解释性的普遍负面发现。在POPE——这三种架构共有的基准测试上——中层在68-91%的错误中编码了真实答案,然而这一信号对最终预测并非因果有效:残差流修补在三种架构的层级别上产生了0%的非平凡翻转,在三种中的两种架构的头级别上也是如此(Qwen:0/12,600次修补前向传播)。唯一的例外是InternVL3的第20层第2个头,这是一个非词汇、自注意的头,其效应局限于该特定头(p < 1e-4)。尽管结果为空,这些错误在操作上可分离为三种失败模式——感知失败、编码但断开、先验覆盖——在三种架构上均可学习到超过60%的准确率,并且架构的先验方向预测了两种干预中哪一种会引发类别特定的响应。我们报告这些在神谕标签下的缓解效应,作为这些类别在机制上真实的证据,而非作为可部署的方法。

英文摘要

Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level (Qwen: 0/12,600 patched forwards). The lone exception, InternVL3 layer-20 head-2, is a non-vocab, self-attending head whose effect is localized to that specific head (p < 1e-4). Despite the null, the errors separate operationally into three failure modes -- Perception Failure, Encoded-but-Disconnected, Prior-Override -- learnable above 60% on all three architectures, and the architecture's prior direction predicts which of two interventions elicits a category-specific response. We report these mitigation effects under oracle labels as evidence the categories are mechanistically real, not as a deployable method.

Comments13 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑