AI 中文总结
该研究提出组合代码编码与N选最优搜索的攻击,突破平均防御成功率99%的SAGE自检查防御,在三个开源目标及70B目标上取得高成功率,并揭示防御借势及攻击排序反转的机制,修复了自身流程的有效性缺陷。
AI 中文摘要
自检查防御要求目标模型在回答前评估请求;已发布的最强实例SAGE报告平均99%的防御成功率。我们展示,可通过组合两种单独对SAGE无害的攻击突破该防御:一种是已有的代码补全编码,另一种是已有的N选最优搜索,二者单独对行为的成功率均不超过4.7%。组合后,将搜索预算用于编码,在三个开源目标上分别达到67%、22%、15%的成功率,且该效果在70B规模的目标上仍持续存在。随后我们对该组合机制进行解释,而非仅报告结果:其一,自检查防御的优势借自目标模型——SAGE本身未检测到攻击,而是要求模型自行判断,四个目标模型将该请求明确拒绝的比例在32%至97%之间,这决定了防御覆盖范围的排序,尽管未防御的可达性几乎一致;其二,哪种攻击能突破取决于防御类型,且排序会反转:针对transform防御,代码编码保留的未防御可达性远高于字符搜索,而针对gate防御则排序反转,我们用攻击向防御决策边界提供的独立探测数量解释这一差异;最后,我们报告在自身流程中发现并修复的一个有效性缺陷:贪心解码下的确定性攻击完全没有N选最优的变异通道,并给出检测该缺陷的一行诊断代码。所有结论均基于经人工验证的评判器评分的310000次生成结果。
英文摘要
Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. We ask what a budget buys when its variance is moved into a structural channel instead, holding the search identical across both arms so the encoding is the only difference. Against SAGE, the strongest published self-check defense, best-of-N over a code-completion encoding reaches 67, 22 and 15% of behaviors on three open-weight targets, where that encoding fired once reaches at most 4.7% and the published character search at full budget at most 3.0%: 9 to 75 times the sum of the parts, with bootstrap intervals clearing both ingredients on every target. We report the operative figure beside the headline rather than the headline alone: at the actionable severity threshold those cells read 24, 8 and 1 behaviors (95% CI [13, 28], [3, 13], [0, 3]). A 2x2 holding encoding and variation apart shows the two defense families fail to different factors: a transform defense is broken by the depth of the encoding (7 -> 67 behaviors at fixed variation) and a gate by the breadth of the variation (13 -> 57 at fixed encoding). Repeated sampling also inflates apparent robustness, because an attacker who may try N times experiences the maximum over draws while safety results are reported as means: on one target SAGE blocks 99.8% of individual draws yet loses 12 behaviors to a repeat attacker where a classifier gate blocking 95.6% loses 10. Removing the target's sampling costs SAGE 59, 76, 82 and 29 points of coverage more than it costs an undefended control, against 25, -5, 8 and 2 for a defense whose verdict comes from a fixed shadow model. The design that loses is the one fusing screening and answering into a single generation, so every attacker draw redraws the safety decision as well. The prescription is architectural, not free: do not fuse screening with generation.
Comments27 pages (7 pages main text, references, 18 pages supplementary material), 2 figures, 22 tables