更安全的内容还是更坚定的拒绝?一种针对有害微调的对齐混合扰动防御
Safer Content or Firmer Refusals? A Hybrid Perturbation Defense for Alignment under Harmful Fine-tuning
浏览论文内容
中文总结 AI 辅助
针对有害微调攻击,提出VaccineBooster混合防御,结合嵌入扰动与梯度衰减,在Llama-2-7B上实现最低有害内容评分0.315,同时揭示内容安全与拒绝保留之间的权衡。
中文摘要 AI 辅助
微调即服务允许用户将安全对齐的语言模型适配到自己的数据上,但这也带来了有害微调的攻击面:在原本良性的微调数据集中混入少量有害数据,即可削弱模型的对齐能力。两种近期的对齐阶段防御方法从模型的不同层面解决这一问题。Vaccine通过提高隐藏嵌入对有害微调引起的表示偏移的鲁棒性来发挥作用,而Booster则模拟有害权重更新并在对齐过程中削弱其影响。我们研究了这两种机制是否具有互补性,并提出了VaccineBooster,一种在每次训练步骤中同时结合嵌入扰动和权重级梯度衰减的单一对齐流程。在使用BeaverTails对齐并随后通过投毒微调攻击的Llama-2-7B上,VaccineBooster在比较的防御方法中取得了最低的OpenAI审核分数0.315,而仅使用Booster的变体保持了最高的攻击后拒绝率50%。结合对嵌入扰动强度和梯度衰减强度的消融实验,这些结果表明存在一种权衡:嵌入扰动主要减少被标记的有害内容,而梯度衰减主要保留明确的拒绝行为。由于我们的评估使用了十个提示且每种配置仅运行一次未设置随机种子,我们将这一权衡报告为观察到的模式而非统计上已解决的效果。这些结果为在对齐模型暴露于不可信微调时,优先考虑内容安全还是拒绝保留提供了实用指导。
英文摘要
Fine-tuning-as-a-service lets users adapt a safety-aligned language model to their own data, but it also creates a harmful fine-tuning attack surface: a small amount of harmful data mixed into an otherwise benign fine-tuning set can degrade the model's alignment. Two recent alignment-stage defenses address this problem at different levels of the model. Vaccine improves the robustness of hidden embeddings to the representation shifts induced by harmful fine-tuning, whereas Booster simulates harmful weight updates and attenuates their effect during alignment. We investigate whether these mechanisms are complementary and propose VaccineBooster, a single alignment procedure that combines embedding perturbation and weight-level gradient attenuation within each training step. On Llama-2-7B aligned with BeaverTails and then attacked through poisoned fine-tuning, VaccineBooster achieves the lowest OpenAI moderation score among the compared defenses, 0.315, while a Booster-Only variant retains the highest post-attack refusal rate, 50%. Together with ablations over the embedding-perturbation and gradient-attenuation strengths, these results indicate a trade-off: embedding perturbation primarily reduces flagged harmful content, whereas gradient attenuation primarily preserves explicit refusal behavior. Because our evaluation uses ten prompts and a single unseeded run per configuration, we report this trade-off as an observed pattern rather than a statistically resolved effect. These results provide practical guidance for prioritizing content safety or refusal retention when aligned models are exposed to untrusted fine-tuning.
发表机构
- University of Louisville(路易斯维尔大学)
机构由 AI 辅助整理,请以论文原文为准。