这个检查点被“消拒”了吗?一种双信号审计及其失败图谱
Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map
浏览论文内容
中文总结 AI 辅助
提出一种无需阈值的双信号审计方法,结合激活拒绝间隙和权重恢复能量,在273个检查点注册表上以AUROC 0.95检测消拒,并映射两种失败模式。
中文摘要 AI 辅助
平台能否在部署前判断一个开放权重检查点的拒绝机制是否被剥离?运行时防护措施无法做到:它们评分的是生成内容,而非工件本身。我们结合两种廉价内部信号——参考锚定的激活拒绝间隙和基础到候选权重差异的权重恢复能量——形成一种无需阈值的检查点审计。两者负相关且标签互补:间隙提供拒绝特异性,权重能量提供召回率。在涵盖Qwen、DeepSeek-distilled Qwen、Llama和Gemma的273个检查点注册表上,它们的z分数和以AUROC 0.95将57个公开消拒与37个良性微调、合并和指令微调区分开,显著高于任一单独信号(0.84、0.90),且Youden校准阈值转移到保留族系时平衡准确率为0.89(假阳性率0.11),仅遗漏57个中的4个。然后我们按严重程度映射两种失败:伪造参考无需训练即可规避两个轴(ΔW=0,ρ=1由构造决定);白盒所有者训练检查点越过阈值,同时保持防护不安全且连贯。该审计是有效的分诊,而非防篡改:它假定一个经过认证的参考,其声明受我们评估的注册表限制。
英文摘要
Can a platform tell, before deployment, whether an open-weight checkpoint has had its refusal mechanism stripped? Runtime guards cannot: they score generations, not the artifact. We combine two cheap internal signals, a reference-anchored activation refusal-gap and a weight-recovery energy of the base-to-candidate weight difference, into a threshold-free checkpoint audit. The two are negatively correlated and label-complementary: the gap supplies refusal-specificity and the weight energy supplies recall. On a 273-checkpoint registry spanning Qwen, DeepSeek-distilled Qwen, Llama, and Gemma, their z-sum separates 57 public abliterations from 37 benign fine-tunes, merges, and instruction-tunes at AUROC 0.95, significantly above either signal alone (0.84, 0.90), and a Youden-calibrated threshold transfers to held-out families at balanced accuracy 0.89 (FPR 0.11), missing only 4 of 57. We then map two failures, in order of severity: a spoofed reference evades both axes with no training (ΔW=0, \r{ho}=1 by construction), and a white-box owner trains a checkpoint past the threshold while it stays guard-unsafe and coherent. The audit is effective triage, not tamper-proofing: it presumes an attested reference, and its claims are bounded by the registry we evaluate it on.
发表机构
- Moonsong Labs(Moonsong实验室)
机构由 AI 辅助整理,请以论文原文为准。