AI 中文总结
研究如何在开放权重模型中防止有害微调并保留良性适应性,提出HarmAlign方法,通过沿估计子空间变形等操作,推导有限样本界,实现条件收敛率控制,在多种设置下有效阻止攻击且保护良性任务。
AI 中文摘要
短时间的微调可能会解除开放权重模型的安全防护,比如重新训练拒绝训练的助手以协助武器开发或产生仇恨言论。在保留良性适应性的同时防止这种有害微调仍然很困难:唯一具有明确曲率证明的先前方法——光谱变形,会全局增加曲率,从而阻碍良性适应和有害适应。我们提出了HarmAlign,它沿着估计的对比激活子空间应用保持函数的光谱变形。我们推导了估计子空间能量和由此产生的局部有害分布曲率下限的有限样本界。恒定步长梯度下降的稳定性——进展二分法将经过验证的曲率转化为条件收敛率控制。从经验上看,在固定架构、有限预算的一阶威胁模型中,HarmAlign在危险知识重新学习设置和有害辅助微调设置中阻止直接微调以及三种数据或目标自适应攻击,同时受保护的良性任务仍然可训练。这种阻止在每个攻击检查点的测试一阶优化器变体中持续存在,并且在分布外有害微调下,它扩展到我们威胁模型中的重要情况:意外安全降级和紧急失调。
英文摘要
A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produce hate speech. Preventing such harmful fine-tuning while retaining benign adaptability remains difficult: the only prior method with an explicit curvature certificate, spectral deformation, inflates curvature globally and thereby obstructs benign adaptation along with harmful adaptation. We propose HarmAlign, which applies function-preserving spectral deformation along a estimated contrastive activation subspace. We derive finite-sample bounds for the estimated subspace energy and the resulting local harmful-distribution curvature lower bound. A stability--progress dichotomy for constant-step gradient descent turns the certified curvature into conditional convergence-rate control. Empirically, within a fixed-architecture, finite-budget first-order threat model, HarmAlign blocks direct fine-tuning and three data- or objective-adaptive attacks across a hazardous-knowledge relearning setting and a harmful-assistance fine-tuning setting, while the protected benign tasks remain trainable. The block persists across the tested first-order optimizer variants over every attack checkpoint, and under out-of-distribution harmful fine-tuning, and it extends to important cases in our threat model: accidental safety degradation and emergent misalignment.
CommentsUnder submission AAAI 2027