发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过SSU-Bench数据集和因果追踪方法,发现视觉-语言模型中联合安全判定受干预影响的位置从早层转向晚层,且最终标记状态可读出模型决策,但可读性不等于正确性。
AI 中文摘要
视觉-语言模型可能需要将图像与提示词结合,以识别单独由任一者都无法揭示的安全风险。这种联合安全判断在模型内部何处变得可访问?我们引入了SSU-Bench,一个由匹配的安全和不安全图像-文本组合构成的数据集,这些组合通过单项提示词编辑或带有标注目标区域的图像编辑构建。使用三个视觉-语言模型,我们在配对输入之间转移内部状态,并测量由此产生的安全判定变化。跨模型及两种反事实类型,在较早的解码器层中,对改变的输入位置进行干预是有效的,而对最终输入标记的干预则在较晚层变得有效。从其他示例估计的方向产生类似的晚层效应。对最终标记状态的线性读出也能预测模型自身的判定,包括错误判断,跨模型比较揭示了反事实变化模式中的相似性。这些发现识别了干预可以影响联合安全判定的位置中反复出现的转变,并将可读的模型决策与正确的安全判断区分开来。
英文摘要
A vision-language model may need to combine an image with a prompt to recognize a safety risk that neither reveals alone. Where does this joint safety judgment become accessible inside the model? We introduce SSU-Bench, a dataset of matched safe and unsafe image-text combinations constructed using single-item prompt edits or image edits with annotated intended regions. Using three vision-language models, we transfer internal states between paired inputs and measure the resulting change in the safety verdict. Across models and both types of counterfactual, interventions at the changed input positions are effective in earlier decoder layers, while interventions at the final input token become effective later. Directions estimated from other examples produce similar late-layer effects. A linear readout of the final-token state also predicts the model's own verdict, including incorrect judgments, and cross-model comparisons reveal similarities in the patterns of counterfactual change. These findings identify a recurring transition in where interventions can influence a joint safety verdict and distinguish a readable model decision from a correct safety judgment.