arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

梯度免疫:对恶意微调的零空间抗性

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

Yuxuan Huang, Xingyu Zeng, Tianhang Zheng, Chaochao Lu

arXiv 2608.05045首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University; Shenzhen University of Advanced Technology; Zhejiang University(上海人工智能实验室; 上海交通大学; 深圳理工大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对部分保护开放权重(PPOW)发布场景的大型语言模型,提出单向安全门(USG)防御,可抑制恶意微调的有害梯度,提升恶意下游适应的成本且无需下游用户配合。

AI 中文摘要

已发布的对齐大型语言模型仍易受恶意下游微调攻击。现有防御措施大多针对微调即服务(FTaaS)范式设计,或依赖下游用户遵循额外安全流程,因此无法直接解决本研究关注的场景:提供商控制的部分保护开放权重(PPOW)发布场景,其中大部分权重保持可训练,仅保留一个小型安全关键组件。我们提出单向安全门(USG),实例化为零空间立方层(Null Space Cubic Layer)与逆适配器(Inverse Adapter),插入在Transformer最后一层之后。在下游微调期间,立方层会抑制或阻止隐藏状态落入校准保护区域的有害样本的梯度,而逆适配器则恢复基础模型的前向行为。实际应用中,防御者使用持有的有害数据校准阈值,使保护能泛化到附近同分布有害样本。在6种评估的模型-数据集设置中,USG在固定发布阈值下将微调后的攻击成功率保持在接近发布前水平,同时在较简单设置中维持高安全通过率,且在BeaverTails的不安全样本上表现出更清晰的安全-效用权衡。这些结果表明,发布时的表示空间阻塞可提高恶意下游适应的成本,无需下游用户配合。代码可在此https URL获取。

英文摘要

Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the fine-tuning-as-a-service (FTaaS) paradigm or rely on downstream users to follow additional safety procedures, and therefore do not directly address the setting we study: a provider controlled partially protected open-weight (PPOW) release setting in which most weights remain trainable while a small safety-critical component is preserved at release. We propose a Unidirectional Safety Gate (USG), instantiated as a Null Space Cubic Layer together with an Inverse Adapter inserted after the final Transformer layer. During downstream fine-tuning, the cubic layer suppresses or blocks gradients from harmful samples whose hidden states fall in a calibrated protected region, while the Inverse Adapter restores the base model's forward behavior. In practice, we calibrate a threshold using defender-held harmful data, allowing protection to generalize to nearby in-distribution harmful samples. Across six evaluated model-dataset settings, USG keeps post-finetuning attack success rate close to the pre-release level under a fixed release threshold, while maintaining high safe-pass rates on easier settings and exhibiting a clearer safety-utility trade-off on unsafe samples from BeaverTails. These results suggest that release-time representation-space blocking can raise the cost of malicious downstream adaptation without requiring downstream cooperation. The code is available at https://github.com/OpenCausaLab/Gradient-Immunity.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑