arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

这是谁的拒绝?黑盒多模态安全护栏的未被衡量的贡献

0%, 45%, or 99%: A Guardrail's Own Share of the Refusals It Is Credited With

Haoyu Zhang, Xiangchen Guan, Yang Chen, Haowen Xu, Shibo Zheng, Xiao Luo, Zhuoxi Wang, Yi Feng, Mohammad Zandsalimy, Shanu Sushmita

arXiv 2608.08641首次发表:更新:

AI 中文总结

该研究指出黑盒多模态安全护栏的评估指标混淆了护栏与目标模型的弃权贡献,通过实验揭示了有效载荷通道和内部读取文本对护栏贡献的影响,并修正了相关评估结果。

AI 中文摘要

黑盒安全护栏的评估仿佛其获得的安全数值完全属于自身,但事实并非如此。受保护的管道包含两个可以执行弃权(不执行)的组件:安全护栏和因自身对齐问题产生弃权的目标模型,所有报告的指标都是两者的总和。我们表明,安全护栏实际获得的安全贡献占比从无到几乎全部,由两项未被任何评估记录的变量决定:承载有效载荷的通道,以及测试工具(harness)在防御内部读取的文本。由于安全护栏会替换模型的响应,且两种计数互不重叠,因此无需额外成本即可恢复该占比。针对两个开放权重目标的文本安全护栏实验显示:当有效载荷以像素形式呈现时,安全护栏不会阻止任何内容,系统产生的所有弃权均由模型生成;当读取攻击者实际发送的编码提示时,安全护栏产生的弃权仅占归因于它的弃权的一小部分;当读取攻击背后的未编码请求时,安全护栏几乎阻止所有内容,模型完全沉默。这种并非源于不精准:相同的安全护栏也不会阻止任何良性图像输入,因此其图像通道的决策是一个常量。授予未编码请求会大幅增加安全护栏的测量收益,对于字幕介导的重新检查收益较少,对于多数投票平滑器则无收益;该排序在独立重复实验中可复现。在一项防御内隔离授予操作显示,其并未提升检测效果:伤害判定阶段无贡献,而重新生成答案的阶段才产生效果。且该被放大的设置并非疏忽:参考实现从单个提示字段构建所有阶段,该字段无法区分攻击者发送的内容与基准记录的内容,因此忠实移植会静默引入该问题。我们此前发表的相关图表也在修正范围内。

英文摘要

A defended pipeline's refusals have two producers: the guardrail bolted in front of the model, and the model's own alignment. Recovering the split costs nothing, because a guard block replaces the model's response and the two counts are therefore disjoint. Holding the defense, the targets, the corpus and the judge fixed, the guardrail's own share of the refusals credited to it is 0%, 41-45%, or 99% across three settings that a results table would describe identically. Two choices move it, and neither belongs to the deployer who bought the guardrail. The attacker drives the share to zero by choosing which channel carries the payload: a plainly written request rendered as pixels, with nothing obfuscated, leaves a text guard's read covering none of it. The evaluator drives the share to 99% by choosing what text fills a defense's internal slots: fill them with the unencoded request behind an encoded attack, a read no deployed defender possesses, and the same guard blocks almost everything. The two consequences differ, and only the attacker's can happen to a running system. The evaluator's choice is an artifact carried by the literature, and its size is set by where the granted text lands: substantial at a guard gate, smaller in a caption-mediated re-check, absent in a majority-vote smoother, an ordering reproduced in an independent replicate. Isolating the grant inside the caption-mediated defense refutes the prediction we registered, since the harm-verdict stage contributes nothing while the stage that regenerates the answer carries the whole effect. The reference implementation builds every stage from a single prompt field that cannot represent the difference between what the attacker sent and what the benchmark records, and an audit of four further released harnesses and of the benchmark itself finds the same structural gap, so faithful porting supplies the grant silently.

Comments34 pages (7 pages main text, references, 24 pages supplementary material), 3 figures, 23 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑