发表机构
University of the Philippines(菲律宾大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FLIP通过最终层推理时探测验证视觉-语言模型中干预位点的机制解释,提出四标准协议,确认面向logit的归一化后状态支持任务关联计算,而非一般扰动。
AI 中文摘要
我们提出了FLIP,一种最终层推理时探测方法,用于测试开放权重视觉-语言模型(VLM)中面向logit的干预位点是否支持结构化的、任务关联的计算,而非一般的扰动。否则,内部干预下的行为变化在机制上是不明确的:它可能反映了视觉证据利用的改善、一般的输出不稳定性或彻底的退化。FLIP在logit计算前对最终归一化的隐藏状态应用逐元素下限操作,保持参数、提示和解码不变。在一个受控的检测/计数探测中,扫描干预强度揭示了三个区域:可忽略的变化、一个有界的内部区域(在该区域内,IoU 0.50($R_{50}$)下的检测召回率提高,而容错计数误差($\mathcal{E}_{\mathrm{count}}$)下降)以及过度抑制。我们形式化了一个四标准探测-扫描协议,用于规范干预效应的解释:区域结构、接地代理对齐、特征一致性依赖,以及未能基于性能的阴性对照上重现相同的正区域。传递给输出头的归一化后状态是此测试的面向logit的实例;在非目标下限扫描下,它满足完整协议。原始解码器层干预,包括最终归一化前的最后块输出,以及单例对左/右对照,未能重现最终位点的特征,而同位点操作符和多个VLM则重现了该特征。因此,FLIP是干预式机制可解释性的验证步骤,而非转向方法。
英文摘要
We present FLIP, a final-layer inference-time probe for testing whether a logit-facing intervention site in an open-weight vision-language model (VLM) supports structured, task-linked computation rather than generic perturbation. Behavioral change under internal intervention is otherwise mechanistically ambiguous: it may reflect improved use of visual evidence, generic output instability, or outright degradation. FLIP applies elementwise flooring to the final normalized hidden state before logit computation, leaving parameters, prompts, and decoding unchanged. On a controlled detection/counting probe, sweeping intervention strength reveals three regions: negligible change, a bounded interior regime in which detection recall at IoU 0.50 ($R_{50}$) improves while tolerant counting error ($\mathcal{E}_{\mathrm{count}}$) falls, and over-suppression. We formalize a four-criterion probe-and-sweep protocol for disciplining the interpretation of intervention effects: regime structure, grounding-proxy alignment, feature-coherence dependence, and failure to reproduce the same positive regime on a performance-based negative control. The post-normalization state passed to the output head is the logit-facing instantiation of this test; under a non-targeted flooring sweep it satisfies the full protocol. Raw decoder-layer interventions, including the last-block output before final normalization, and the singleton-pair left/right control fail to reproduce the Final-site signature, while same-site operators and multiple VLMs replicate it. FLIP is therefore a validation step for intervention-based mechanistic interpretability, not a steering method. Repo at https://github.com/earl-juanico/flip-qwen3vl/
Comments25 pages, 14 figures, 5 tables. Accepted at the Mechanistic Interpretability Workshop at ICML 2026, Seoul, South Korea