arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.18931cs.LGcs.AI

Safe Pruning LoRA:针对LLMs适配中安全对齐的鲁棒距离引导剪枝

Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs

  • School of Electronics and Computer Science, University of Southampton, UK(索姆塞特大学电子与计算机科学学院)
  • Department of Computer Science, University of Liverpool, UK(利弗pool大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

Shuang Ao, Yi Dong, Jinwei Hu, Sarvapali Ramchurn

更新

AI总结:

本文提出SPLoRA框架,通过引入E-DIEM度量检测安全错位并剪枝削弱安全对齐的LoRA层,在保持或提升模型性能与可靠性的同时显著降低安全风险并减少推理开销。

AI中文摘要:

使用Low-Rank Adaptation (LoRA)对Large Language Models (LLMs)进行微调在增强适应性的同时降低了计算成本。然而,即使使用良性数据,微调也可能破坏安全对齐,增加产生有害输出的风险。现有的安全对齐方法难以捕捉复杂的参数偏移,导致安全性与实用性之间的权衡不佳。为解决此问题,我们提出了Safe Pruning LoRA (SPLoRA),这是一种新颖的基于剪枝的方法,可选择性地移除削弱安全对齐的LoRA层,在保持性能的同时提升安全性。其核心是我们引入了Empirical-DIEM (E-DIEM),这是一种对维度不敏感的相似度度量,能有效检测LoRA适配模型中的安全错位。我们在混合良性与恶意数据以及纯良性数据集微调的LLMs上进行了广泛实验,从实用性、安全性和可靠性指标对SPLoRA进行了评估。结果表明,SPLoRA优于最先进的安全对齐技术,在维持或提升模型性能与可靠性的同时,显著降低了安全风险。此外,SPLoRA减少了推理开销,使其成为部署更安全、更可靠LLMs的可扩展且高效的解决方案。代码可在https://github.com/AoShuang92/SPLoRA获取。

英文摘要:

Fine-tuning Large Language Models (LLMs) with Low-Rank Adaptation (LoRA) enhances adaptability while reducing computational costs. However, fine-tuning can compromise safety alignment, even with benign data, increasing susceptibility to harmful outputs. Existing safety alignment methods struggle to capture complex parameter shifts, leading to suboptimal safety-utility trade-offs. To address this issue, we propose Safe Pruning LoRA (SPLoRA), a novel pruning-based approach that selectively removes LoRA layers that weaken safety alignment, improving safety while preserving performance. At its core, we introduce Empirical-DIEM (E-DIEM), a dimension-insensitive similarity metric that effectively detects safety misalignment in LoRA-adapted models. We conduct extensive experiments on LLMs fine-tuned with mixed of benign and malicious data, and purely benign datasets, evaluating SPLoRA across utility, safety, and reliability metrics. Results demonstrate that SPLoRA outperforms state-of-the-art safety alignment techniques, significantly reducing safety risks while maintaining or improving model performance and reliability. Additionally, SPLoRA reduces inference overhead, making it a scalable and efficient solution for deploying safer and more reliable LLMs. The code is available at https://github.com/AoShuang92/SPLoRA.

补充信息

↑