arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08236cs.AI

风格胜于实质:内容不变包装器翻转LLM安全裁判的判定

Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts

Yongxi Zhou, Wenbo Ye, Yuanzhe Liu, Zihan Dong, Junwei Yao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过内容不变的风格包装器测试LLM安全裁判,发现特定裁判存在可被廉价利用的盲点,如令牌拒绝包装器翻转GPT-4o-mini近20%的正确判定,而Llama Guard 4可被确定性欺骗,漏洞在于裁判而非内容。

中文摘要 AI 辅助

自动安全裁判——诸如Llama Guard或GPT-4o评分提示等系统,用于判断模型回复是否有害——为几乎所有已报告的越狱成功率、防御评估和安全排行榜提供了数据。我们探究这些裁判评判的是回复的内容还是其表达方式。我们保持回复内容不变,添加内容不变的风格包装器:在回复之前或之后放置的固定字符串,仅改变其语气(如教育性免责声明、虚假的安全“推理”块、先拒绝后跟未改变的有害主体),或者对于无害的拒绝,添加仅听起来危险的框架。主体内容逐字节保留,因此忠实的裁判必须返回相同的判定,任何翻转都是裁判的错误,而非安全性的变化。在超过600条JailbreakBench回复×最多7种形式×8个裁判的实验中,我们通过配对显著性检验和测量的噪声底限来测量翻转率。发现是精确而非普遍的:大多数裁判几乎不变,但特定裁判存在可廉价利用的盲点。一个令牌拒绝包装器翻转了GPT-4o-mini 19.9%的正确“不安全”判定(噪声底限0.5%;在三人多数重新评分下为18.2%),而仅使Claude变化0.4%。已部署的Llama Guard 4被确定性玩弄:“教育课程”框架将其12.3%的有害判定翻转为安全。第二个已部署的防护(gpt-oss-safeguard-20b)免疫,仅重写评分提示(StrongREJECT风格)在相同模型上将攻击削减了十倍——漏洞存在于裁判而非内容中。双注释者人工验证确认了100%的内容不变性和90%的翻转属于裁判错误(kappa 0.95-1.0),自举法显示底层模型排名仅因采样就已不稳定。我们发布了数据集、包装器、代码和每个判定的标签。

英文摘要

Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model's reply is harmful -- produce the numbers behind almost every reported jailbreak success rate, defense evaluation, and safety leaderboard. We ask whether these judges grade what a reply contains or how it sounds. We keep a reply's content fixed and add content-invariant style wrappers: fixed strings placed before or after the reply that change only its tone (an educational disclaimer, a fake safety "reasoning" block, a token refusal followed by the unchanged harmful body), or, on harmless refusals, framing that merely sounds dangerous. The body is preserved byte-for-byte, so a faithful judge must return the same verdict, and any flip is an error of the judge, not a change in safety. Over 600 JailbreakBench replies x up to 7 forms x 8 judges, we measure flip rates with paired significance tests and measured noise floors. Findings are precise rather than universal: most judges barely move, but specific judges harbor cheaply exploitable blind spots. A token-refusal wrapper flips 19.9% of GPT-4o-mini's correct "unsafe" verdicts (noise floor 0.5%; 18.2% under majority-of-three re-scoring) yet moves Claude only 0.4%. The deployed Llama Guard 4 is deterministically gamed: an "educational course" framing flips 12.3% of its harmful verdicts to safe. A second deployed guard (gpt-oss-safeguard-20b) is immune, and rewriting only the grading prompt (StrongREJECT-style) cuts the attack tenfold on the identical model -- the vulnerability lives in the judge, not the content. A two-annotator human validation confirms 100% content invariance and 90% of flips as judge errors (kappa 0.95-1.0), and a bootstrap shows the underlying model ranking is already unstable to sampling alone. We release the dataset, wrappers, code, and per-verdict labels.

发表机构

  • Northeastern University(东北大学)
  • University of Southern California(南加州大学)
  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑