arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

拒绝局部化,损伤转移:少样本微调下的安全层

Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Heuiseok Lim

arXiv 2610.00320首次发表:更新:

发表机构

Korea University; Zoom Communications; Yonsei University Mirae Campus(高丽大学; Zoom 通信公司; 延世大学未来校区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现微调攻击可移除大模型拒绝能力,基于定位的修复在攻击变化下失效,需自适应防御检查。

AI 中文摘要

微调使对齐的大型语言模型(LLMs)适应下游任务,但几十个有害示例即可移除其对有害请求的拒绝。先前的工作将安全相关行为定位到特定层、方向和标记,提出了保护目标。我们测试成功的定位与恢复是否支持在攻击变化后仍存活的防御。在来自四个模型家族的六个检查点上,有害与良性提示在攻击后仍保持线性可分,将完整的干净隐藏状态修补到受损模型可在可复现的转变深度恢复拒绝。基于先前的层冻结防御,我们冻结直至该深度的每一层并重复攻击。在100个有害示例下,所有六个检查点的拒绝率仍接近零,恢复转变位于冻结边界之上。在第二项研究中,在四个检查点上进行仅注意力的短微调后,移除更新的前两个奇异方向可恢复拒绝。在Llama-3.1-8B上,普通训练变化削弱了此修复,而将更新分散的攻击者则能击败它。在良性Llama微调上校准的谱检测器在该检查点上遗漏了大多数修复失败。然而,局部冻结在少量有害示例无意进入训练数据时仍有助于保持拒绝。这些结果表明,攻击者可以绕过由恢复识别的区域,并击败在多个检查点上有效的修复,这促使对自适应微调防御提出五项检查。代码可在以下https URL获取。

英文摘要

Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization and recovery support defenses that survive changes in the attack. Across six checkpoints from four model families, harmful and benign prompts remain linearly separable after attack, and patching full clean hidden states into the compromised model restores refusal at a reproducible transition depth. Building on a prior layer-freezing defense, we freeze every layer up to this depth and repeat the attack. At a hundred harmful examples, refusal remains near zero on all six checkpoints, with recovery transitions above the frozen boundary. In a second study, removing the update's top two singular directions restores refusal after short attention-only fine-tunes on four checkpoints. On Llama-3.1-8B, ordinary training changes weaken this repair and an attacker who spreads the update defeats it. A spectral detector calibrated on benign Llama fine-tunes misses most repair failures on that checkpoint. Localized freezing can nevertheless help preserve refusal when a few harmful examples enter training data unintentionally. These results show that an attacker can bypass a region identified by recovery and defeat a repair that works across multiple checkpoints, motivating five checks for defenses against adaptive fine-tuning. Code is available at https://github.com/js-lee-AI/refusal-relocates.

Comments24 pages, 7 figures, 21 tables. Jungseob Lee and Dongyub Jude Lee contributed equally

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑