用于鲁棒即插即用适应的解耦对齐
Decoupled Alignment for Robust Plug-and-Play Adaptation
- Northwestern University(西北大学)
- New York University Abu Dhabi(纽约大学阿布扎克分校)
- Stanford University(斯坦福大学)
- Texas A&M University(德克萨斯农工大学)
- University of California, Los Angeles(加州大学洛杉矶分校)
- Illinois Institute of Technology(伊利诺伊理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究提出无需训练的大语言模型对齐方法,利用知识蒸馏提取对齐信号,经模型融合实现即插即用的对齐校正,采用增量调试识别关键知识组件,在有害问题数据集上显著提升防御成功率,且不损性能。
AI中文摘要:
我们引入一种无需训练的安全增强方法,用于对齐大语言模型,无需监督微调或从人类反馈进行强化学习。主要思想是提供一种鲁棒的即插即用方法,防止模型适应下游任务时的影子对齐。具体利用知识蒸馏从对齐良好的大语言模型中提取对齐信号,并通过模型融合注入到影子对齐模型中,实现即插即用的对齐校正。采用增量调试识别有效蒸馏所需知识的关键组件。在有害问题数据集上,该方法显著提高平均防御成功率约14.42%,在17个受影响的大语言模型上高达51.39%,且不影响性能。代码可通过链接获取。
英文摘要:
We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback. Our main idea is to provide a robust plug-and-play approach to prevent shadow alignment when models are adapted to downstream tasks. Specifically, we leverage knowledge distillation to extract alignment signals from well-aligned LLMs and inject them into shadow-aligned models via model fusion, enabling plug-and-play alignment correction. In our methodology, we employ delta debugging to identify the critical components of knowledge necessary for effective distillation. On the harmful question dataset, our method significantly enhances the average defense success rate by approximately 14.42%, reaching as high as 51.39% across 17 influenced LLMs, without compromising performance. Our code is available at https://github.com/NWULIST/DAPA.