AI 中文总结
研究发现安全对齐的VLM会因是否附加图像(如空白画布)而大幅改变拒答率,且此行为与风险无关、不受指令控制,是随权重继承的不可控变量。
AI 中文摘要
视觉语言模型(VLM)的安全性预期取决于请求所要求的内容。我们表明,经过安全对齐的VLM还会根据请求形式的一个属性来决定是否拒答:是否附加了图像,而保持请求所要求的一切不变。附加一张空白画布——不可读、与请求无关、在所有提示中相同——会使拒答率变化数十个百分点,且循环中没有防御措施。这种变化并非全面的谨慎。中性指令几乎不受影响,而边缘良性提示则剧烈变化,因此代价落在与敏感性相邻的流量上:关于隐私、自残和暴力的良性问题。仅附加图像就足够了,而图像的属性决定了代价:相同尺寸的黑色画布比白色画布代价高得多,在开放检查点上,承载轴是像素数量。这种变化也不受指令控制——告诉模型该图像是应忽略的占位符只能消除其中一小部分,而在一个模型上,断言存在附件会在没有任何附件的情况下显著改变拒答率。在部署中,附件可能与风险相关;但这些模型对其的处理并不追踪风险。这不是服务堆栈的问题,因为通过两种方式到达的相同权重表现相同,也不是VLM本身的属性,因为几个开放权重检查点没有显示任何变化。它属于特定的对齐检查点,其中一个是开放的。它也与它所购买的东西脱钩:画布确实在匹配的有害集上阻止了一些攻击成功,但远低于其代价,且其符号不固定——在一个开放模型上,相同的画布使模型明显更容易被攻击。图像存在性不是部署者选择或定价的默认设置;它是随权重继承的不可控变量。
英文摘要
We show that aligned vision-language models also condition refusal on a property of a request's form: whether an image is attached, holding everything the request asks fixed. Attaching a blank canvas, an image that cannot be read, cannot relate to the request, and is byte-identical across every prompt in its condition, shifts benign refusal by tens of points. The shift is not blanket caution but a threshold shift: genuinely neutral instructions are almost unaffected (<=2 percentage points on three of four hosted models) while borderline-benign prompts move +23 to +51 points, so the cost falls on sensitivity-adjacent traffic, meaning benign questions about privacy, self-harm, violence and illegal activity. There is a benign reading of such a threshold, namely that attachment correlates with risk in real traffic, and we take it seriously; a black-box study cannot measure that correlation and we do not claim to. What it can test is whether the response to attachment is robust to variation carrying no information about the request, and on four independent measurements it is not. It varies with canvas colour and pixel count. Its sign inverts across checkpoints. It survives an explicit instruction to disregard the image. And on one model it fires on a bare assertion that an attachment exists, with nothing attached and the modality word contributing none of it. Finally we price it. On a matched harmful set the same canvas does lower attack success, so the cue buys something. But the charge is decoupled from the purchase: the checkpoint with the least harmful headroom we measure, completing only 2% of plain harmful requests, still pays the benign cost in full, and across our models the harmful-side denominator falls as alignment improves while the benign cost does not track it down. Image presence is not a conservative default that a deployer chose and priced. It is an uncontrolled variable.
Comments27 pages (8 pages main text, references, 18 pages supplementary material), 1 figure, 23 tables