AI 中文总结
针对大语言模型的删除(Abliteration)攻击,提出AMRA权重编辑防御方法,在Llama-3-8B和Gemma-2-9B上有效提升删除后的拒绝能力,同时控制了性能损失。
AI 中文摘要
删除(Abliteration)是一种通过将权重矩阵投影到提取的拒绝方向正交空间来移除大语言模型拒绝能力的攻击,仅用少量对比提示就能绕过后训练对齐,已成为突出的安全问题。现有防御常忽视删除的成因,即拒绝方向的易提取性。为阻碍这一过程,我们提出一种权重编辑方法:对残差流写入矩阵应用秩-$k$更新以模糊拒绝信号,将诱导拒绝的激活替换为随机别名,并修正下游读取矩阵以保留模型原行为。在Llama-3-8B上,AMRA使删除后的拒绝得分较未防御基线提升2.16个百分点,MMLU下降不足0.5个百分点;在Gemma-2-9B上,其使删除后的拒绝得分较基线提升14.70个百分点,有害输出率与基线相近,仅存在更大的效用损失。
英文摘要
Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass post-training alignment using only a small set of contrastive prompts. We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted. To hinder this process, we introduce a weight-editing method that obscures the refusal signal by applying rank-$k$ updates to residual stream writer matrices while replacing refusal-inducing activations with random aliases and correcting downstream reader matrices to preserve the model's original behavior. On Llama-3-8B, AMRA improves post-abliteration refusal scores by $2.16$ points over the undefended baseline with less than $0.5$ percentage points of MMLU degradation. On Gemma-2-9B, it improves the post-abliteration refusal by $14.70$ points over the baseline while keeping harmful output rates similar to the baseline, albeit at a greater utility cost.