arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16204cs.LGcs.CLcs.CR

诱饵方向优化:针对LLM消融的后置防御

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

Aashiq Muhamed, Mona T. Diab, Virginia Smith

首次发表
浏览论文内容

中文总结 AI 辅助

提出诱饵方向优化(DDO),一种无需微调的后置权重编辑防御,通过注入非线性诱饵信号破坏消融攻击的估计器,在多个模型上显著降低攻击成功率且成本极低。

中文摘要 AI 辅助

开放权重语言模型中的安全护栏可被拒绝特征消融(Refusal Feature Ablation, RFA)轻易绕过,该技术从残差流中识别并投影出一个线性拒绝方向,通常在保持模型能力的同时实现高攻击成功率(ASR)。针对这些攻击的防御通常需要对每个新检查点进行计算成本高昂的安全微调。我们提出诱饵方向优化(Decoy Direction Optimization, DDO),一种快速、后置的权重编辑防御方法,无需对基础模型进行微调。我们的方法基于一个简单的机制性见解:消融攻击依赖对比估计器来寻找拒绝方向。DDO并非试图隐藏真实的拒绝电路,而是主动向网络的MLP神经元注入一个高幅度、非线性的诱饵信号。当攻击者试图定位拒绝方向时,诱饵会破坏其估计器,诱骗其消融一个无害的正交特征,而实际的安全机制保持完好。我们证明了一个谱界来形式化这一效应,并在六个模型家族上评估DDO,在标准RFA下实现了低于10%的ASR。在Llama-3-8B-Instruct上,DDO在自适应多阶段攻击下与训练式防御相当(最坏情况ASR为65%对比58%),并将Heretic权重级攻击ASR从88.7%降至18%,且每个配置的优化成本比训练式基线低30至450倍。

英文摘要

Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal direction from the residual stream, often achieving a high attack success rate (ASR) while preserving model capability. Defending against these attacks typically requires computationally expensive safety finetuning for every new checkpoint. We introduce Decoy Direction Optimization (DDO), a fast, post-hoc weight-editing defense that requires no base-model finetuning. Our approach is based on a simple mechanistic insight: ablation attacks rely on contrastive estimators to find the refusal direction. Rather than trying to hide the true refusal circuitry, DDO actively injects a high-magnitude, nonlinear decoy signal into the network's MLP neurons. When an attacker attempts to locate the refusal direction, the decoy corrupts their estimator, tricking them into ablating a harmless orthogonal feature while the actual safety mechanism remains intact. We prove a spectral bound formalizing this effect and evaluate DDO across six model families, achieving <10% ASR under standard RFA. On Llama-3-8B-Instruct, DDO remains comparable to trained defenses under adaptive multi-phase attacks (65% vs. 58% worst-case ASR) and reduces Heretic weight-level attack ASR from 88.7% to 18%, all at 30 to 450 times lower optimization cost per configuration than the trained baselines.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑