发表机构
Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示编码提示攻击的有害分支指标无法反映模型真实安全性,编码破坏的是区分能力而非拒绝能力,且后训练无法修复,需引入对照实验。
AI 中文摘要
编码提示攻击几乎完全在其有害分支上进行评估:基准测试发送混淆的有害请求,并报告模型服从的频率。我们表明,该分支几乎不携带关于被测模型的任何信息。在四个独立后训练的7-8B模型中,对有害同形字编码提示的拒绝率跨度仅为0.08——这落在仅由采样噪声在n=100时产生的0.10上限之内——而相同的四个模型在相同请求的明文形式下拒绝率跨度达0.57。编码所破坏的不是拒绝,而是区分能力:在一个模型上,有害与良性请求拒绝率之间的差距从明文下的+0.82降至编码下的恰好0.00,良性请求和有害请求以相同的0.99被拒绝。仅读取有害分支的基准测试将该模型与另一个保持+0.61差距的模型评分相同。我们随后探究后训练是否能修复这一问题,使用基于相同基础权重的已发表配方。答案是否定的:在完整的SFT -> DPO -> RLVR流程中,明文下的有害区分度从+0.55提升至+0.80,而编码引起的损失保持不变,为0.34-0.50,并且在测试的每种编码下,标准的有害分支指标与区分度的变化方向相反。如果没有该领域不常规运行的对照实验,这些现象均不可见。我们报告了八个仪器缺陷,每个缺陷都附有捕获它的对照实验;它们具有共同的方向,即行为轴上的每个缺陷都夸大了表观安全性。
英文摘要
Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied. A high refusal rate there is reported as safety, and it is equally consistent with a model that has stopped telling the request apart from anything else in the same format. We run the benign arm through the same transformation, and the two cases are far apart. Across four 7-8B models spanning three base families and four post-training recipes, refusal of harmful homoglyph-encoded prompts spans 0.08 while the same four span 0.57 on the identical requests in plaintext. What the encoding destroys is not refusal but the harm gap: on one model the gap between harmful and benign refusal falls from +0.82 in plaintext to exactly 0.00 under the encoding, and a benchmark reading only the harmful arm scores that model and one retaining a +0.61 gap identically. Running the cell such benchmarks leave out (plaintext content wearing the attack template, with nothing obfuscated) shows that on two of the four models the loss is caused by the protocol rather than by the character transformation, and on a third by the characters. Across a full SFT -> DPO -> RLVR pipeline the harm gap rises by +0.26 with a paired interval excluding zero while the standard harmful-arm metric registers no resolved change at all. We report twelve instrument defects, each with the control that caught it, including a binary jailbreak judge that fires on 0.61-0.70 of responses to plaintext benign prompts; six of the twelve inflate apparent safety, which is the direction a broken safety evaluation fails in by default.
Comments16 pages, 1 figure, 7 tables; supplementary material included as an appendix