arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

未读取还是未执行?区分内容防护中的表征失败与执行失败

Unread or Unenforced? Separating Representation from Enforcement Failure in Content Guards

Haoyu Zhang, Yi Feng, Shibo Zheng, Hanwen Liu, Haowen Xu, Xiao Luo, Zhuoxi Wang, Mohammad Zandsalimy, Shanu Sushmita

arXiv 2609.26178首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过残差流探针与判定逻辑分离内容防护的表征失败与执行失败,发现真实策略失败范围比未受控分析更窄,且防护对编码外观无反应。

AI 中文摘要

当编码攻击通过内容防护时,防护要么从未表征载荷的有害内容,要么表征了但未能采取行动。端到端攻击成功率对两者只报告一个数字,然而补救措施是相反的:一个是表征限制,更多安全训练无法触及;另一个是决策规则,安全训练可以触及。我们通过读取防护自身的残差流——一个在明文上拟合且无需重新拟合即可转移到编码条件下的内容探针——连同其判定逻辑(verdict logits)来区分两者,代价仅为一次前向传播且无需评判模型。诚实地做到这一点是问题的大部分,也是我们的主要贡献。置换检验允许在19个条件中的17个上对一个开放防护进行解码测量,对另一个防护则在19个条件中的12个上进行;长度匹配的零分布和基于防护基础模型可证明无法解码的条件校准的控制下限将两者均减少至4个。被丢弃的单元并非边缘单元:我们首次分析中最大的结果——一个防护以AUROC 0.72表征密码但未阻止任何密码——是其基础模型以零速率解码的编码上的伪影。在一个防护上,两个共享无输入的屏幕在拒绝哪些条件上完全一致。幸存下来的是一个真实但比未受控分析声称的更窄的策略失败:在防护大量阻止的条件下,每100个提示中有7至23个被表征且未被阻止,在一个防护几乎不阻止的条件下每100个中有56个。未解码即阻止的情况在整个过程中接近零,因此两个防护都不是对编码出现作出反应,而是对内容作出反应。对于真正的密码,两个防护基本上什么都不阻止,我们将这些单元报告为未测量,而非作为解码失败的证据。

英文摘要

When an encoded attack passes a content guard, the guard either never represented the payload's harmful content or represented it and failed to act. End-to-end attack success rate reports one number for both, yet the two have opposite remedies: one is a representational limit that more safety training cannot reach, the other is a decision rule that it can. We separate them by reading a guard's own residual stream, using a content probe fitted on plaintext and transferred without refitting to the encoded condition, alongside the verdict logits from the same pass. Licensing that read honestly is most of the problem and is our main contribution. A conventional permutation test admits the decode measurement on most of a 19-condition encoding ladder for each of two open guards. A length-matched null and a floor calibrated on conditions the guard's base model provably cannot decode reduce it to four conditions each; holding out the items the probe was fitted on removes one more. A third screen constrains the block axis, which the decode screens leave untouched, by running plaintext content inside each condition's own wrapper. It removes the largest cell that survived them. What remains is a policy failure that survives an item-level holdout on two of the four surviving conditions, at 8 and 7 per 100 prompts, against 17 and 23 when the probe is allowed to have seen the prompt it is scoring. It is also confined to one family of surface encodings: where an encoding leaves content linearly recoverable we can separate the two failures, and on genuine ciphers we report the cells as unmeasured rather than as evidence that nothing was decoded. Across every guard and condition pair, blocked without decoding is near zero, so we find little evidence for a pure encoding-format detector under the conditions we test. That cell is the one read we do not repeat under the holdout, and we report it as such.

Comments14 pages (8 main paper including references, 6 supplementary material), 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑