愚人金:针对开放权重模型安全移除攻击的防御性欺骗
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
- Microsoft Azure(微软Azure)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对开放权重模型易被移除安全对齐的问题,提出“愚人金”防御,通过训练仅在被攻击状态显现的伪造诱饵,使被攻击模型对危险请求输出高比例假回答,且无法区分真假,同时不影响干净状态的良性行为。
中文摘要 AI 辅助
开放权重语言模型的安全对齐极易被移除:通过abliteration(一种将拒绝调解方向从权重中投影的技术)可在数分钟内完成,且目前尚无任何发布时的防御措施能长期阻止该攻击。无法阻止的攻击可被欺骗。我们的防御方案为诱饵加固(“愚人金”),它允许拒绝被移除,但会破坏其收益:一旦拒绝被移除,对危险操作请求的多数回答将是自信、流畅的诱饵,其关键要素均为伪造。诱饵在攻击的可微模拟环境中训练,仅在被攻击状态下显现;拒绝锚点与良性约束则维持干净状态下的原始行为。我们在5个系列的7个模型(规模9B至122B,含稠密模型与混合专家模型)上实现了该方案。在通过预注册效能门槛的6个模型中,对预留提示的被攻击状态响应有0.51至0.90为诱饵,其中0.27至0.84的比例归因于该防御;所有6个模型均符合注册的良性行为与能力预算;第7个(较小模型)未通过门槛(边界情况)。该比例在冻结测试拆分或未触及分层数据中可复现。该主张属于认知层面:若无独立真值,我们测试的所有观测表面均无法区分伪造回答与正确回答——在外部红队基准的CBRNE相关切片上,受防御的122B模型在匹配质量回答中有0.82至0.86为致命错误,而未受防御模型最多为0.10。重复采样无法恢复信任:在工具验证的提示中,K=64的元素级共识仅能重建0.083至0.625的完全可用流程,而未受防御模型为0.58至0.96,且无无标签方式区分两种状态;在最弱模型上,该主张仅适用于单次采样。我们评估了化学与生物危险;该防御不解决上下文越狱问题,仅保护初始发布的受防御权重。
英文摘要
Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening ("Fool's Gold"), concedes the refusal strip and poisons its payoff: once refusal is stripped, most answers to hazardous operational requests are confident, fluent decoys whose critical elements are falsified. Decoys are trained inside a differentiable simulation of the attack, expressing only in the attacked state; a refusal pin and benign leash hold clean-state behavior to the original. We instantiate it on seven models from five families (9B-122B, dense and mixture-of-experts). On the six models passing our pre-registered efficacy gate, 0.51-0.90 of attacked-state responses to held-out prompts are decoys, +0.27-0.84 attributable to the defense; all six stay within registered benign-behavior and capability budgets; the seventh (smaller) fails the gate (boundary case). Rates replicate on a frozen test split or untouched strata. The claim is epistemic: without independent ground truth, no observation surface we tested separates falsified answers from correct ones - on external red-team benchmarks' CBRNE-adjacent slice, the defended 122B is fatally wrong on 0.82-0.86 of matched-quality answers vs at most 0.10 undefended. Repeated sampling does not restore trust: element-wise consensus at K=64 reconstructs a fully usable procedure on 0.083-0.625 of prompts where the instrument validates, vs 0.58-0.96 undefended, with no label-free way to tell the regimes apart; on the weakest such model the claim is per-draw only. We evaluate chemical and biological hazards; the defense does not address in-context jailbreaks and protects only the initially released defended weights.