发表机构
North China Electric Power University; Southeast University; Soochow University(华北电力大学; 东南大学; 苏州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对恶意微调侵蚀大模型拒绝行为的问题,提出SLDR防御方法,通过选择性层恢复和动态路由,在保持下游性能的同时大幅降低有害输出。
AI 中文摘要
微调即服务使用户能够将对齐的大型语言模型(LLM)适配到专门任务,但恶意微调可能侵蚀拒绝行为,同时保留对合法输入的任务性能。我们重新审视了最近的逐层安全诊断,发现安全敏感性是有符号的:缩放不同层可以增强拒绝、削弱拒绝或几乎没有影响。受此观察启发,我们提出了SLDR,一种基于选择性层恢复与动态路由的微调后防御方法。SLDR仅在有符号谱中具有最大和最小敏感性分数的层上训练LoRA恢复适配器,并使用基于表示的动态路由推理,仅对恶意查询激活适配器。在四种模型架构、五个下游任务和四个有害基准上,SLDR大幅减少有害输出,同时保持下游效用。在Llama3.1/SST2上,SLDR将平均有害分数从11.54降至0.08,同时保持下游准确率,且在投毒比例高达0.9时有害分数仍接近零。代码可在该https URL获取。
英文摘要
Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling different layers can strengthen refusal, weaken it, or have little effect. Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing. SLDR trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing inference to activate the adapter only for malicious queries. Across four model architectures, five downstream tasks, and four harmful benchmarks, SLDR substantially reduces harmful outputs while preserving downstream utility. On Llama3.1/SST2, SLDR reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, and the harmful score remains near zero under poisoning ratios up to 0.9. The code is available at https://github.com/Stardust457/SLDR.
CommentsAccepted at NeurIPS 2026