arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HyperSafe:微调语言模型推理时的安全恢复

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

Aznaur Aliev, Carlos Hinojosa, Abdelrahman Eldesokey, Bang An, Bernard Ghanem, Yibo Yang

arXiv 2607.11475首次发表:更新:

发表机构

King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对微调语言模型推理时安全对齐脆弱的问题,提出HyperSafe框架,通过生成特定模型的安全侧网络,利用层激活指纹和超网络映射参数,实现提示级安全分类,有效降低有害响应率,保持任务准确率。

AI 中文摘要

大语言模型中的安全对齐在微调时可能很脆弱,因为即使是良性任务适应也可能增加有害合规性。现有防御主要有两个方向,但存在成本高、损害任务性能或错过特定微调检查点故障等局限。为此提出HyperSafe框架,通过为每个微调检查点生成特定模型的安全侧网络(SSN)来恢复安全行为。利用层激活指纹,通过超网络将其映射到SSN参数,SSN与冻结的微调模型并行运行进行提示级安全分类。在两个模型家族上评估,HyperSafe将有害响应率从19%-31%降至1%以下,同时保持下游任务准确率在微调基线的1%以内。

英文摘要

Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint. These limitations motivate a post hoc, model-specific, and non-invasive approach to safety restoration. To meet these requirements, we propose HyperSafe, a framework that restores safety behavior by generating a model-specific Safe Side Network (SSN) for each fine-tuned checkpoint. HyperSafe uses layer-wise activation fingerprints to capture how fine-tuning changes the model's inner representations. With a small set of given calibration prompts, the hypernetwork maps these fingerprints to the parameters of the \ssn{} in a single forward pass. The generated \ssn{} runs alongside the frozen fine-tuned model and performs prompt-level safety classification: harmful prompts are routed to refusal, while safe prompts are answered by the original fine-tuned model. Thus, HyperSafe requires no gradient updates, no safety data at deployment time, and no modification to the deployed model weights. We evaluate HyperSafe on two model families, Qwen2-7B and LLaMA-3-8B, across multiple safety benchmarks. HyperSafe reduces harmful response rates from 19-31% to below 1% on every held-out checkpoint, while keeping downstream task accuracy within 1% of the fine-tuned baseline on average. Code is available at https://github.com/nokronim/project-safety-remedy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑