弃权几何反映弃权训练:多样的弃权前缀可提升稳定秩并削弱弃权向量消融攻击
Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks
浏览论文内容
中文总结 AI 辅助
本研究以OLMo-2-0425-1B-Instruct为对象,揭示弃权几何与弃权训练的关联,发现多样弃权起始可提升梯度稳定秩,削弱弃权向量消融攻击,为提升AI模型弃权鲁棒性提供新视角。
中文摘要 AI 辅助
弃权训练通过训练模型拒绝不安全查询来保护AI模型免受越狱攻击,降低滥用风险。近期研究发现,对齐语言模型中的弃权行为可由单一激活方向或跨有害提示共享的低维弃权子空间介导:消融这些方向会抑制弃权,同时在很大程度上保留模型的其他能力。然而,目前仍不清楚为何众多模型中的安全关键特征会呈现并集中于低维结构。在对OLMo-2-0425-1B-Instruct的案例研究中,我们发现弃权几何反映了弃权训练:由弃权补全首令牌损失产生的激活更新可解释最终的弃权方向和弃权子空间。我们通过跨弃权数据集的训练动态研究弃权方向,发现其脆弱性与重复的弃权起始相关,而这又与梯度和弃权特征在低维子空间中的集中有关。通过冻结模型分析和受控合成微调,我们发现了一个强化机制:多样的弃权起始可提升梯度和激活变化的稳定秩,使弃权更难被向量消融攻击移除。
英文摘要
Refusal training protects AI models from jailbreaks by training models to decline unsafe queries, reducing the risk of misuse. Recent work finds that refusal behavior in aligned language models can be mediated by a single activation direction or a low-dimensional refusal subspace shared across harmful prompts: ablating those directions suppresses refusals while largely preserves other model capabilities. Yet it remains unclear why safety-critical features in a wide range of models emerge in a concentrated, low-dimensional structure. In a case study of OLMo-2-0425-1B-Instruct we find that the refusal geometry reflects refusal training: activation updates resulting from refusal-completion first-token losses explain the resulting refusal direction and refusal subspace. We study refusal directions through the training dynamics across refusal datasets and reveal that their brittleness is associated with repetitive refusal starts, which in turn is linked to concentration of gradients and refusal features in a low-dimensional subspace. Across frozen-model analyses and controlled synthetic fine-tuning, we find evidence of a hardening lever: diverse refusal starts can raise stable ranks of gradients and activation changes, making refusals harder to remove with a vector ablation attack.
发表机构
- UC San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。