发表机构
LERo(LERo)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对FTaaS平台削弱模型安全对齐的问题,提出TRACE框架,通过模拟有害轨迹生成损坏状态,优化插件补丁,在保持效用的同时恢复模型安全,在多基准测试和模型上占据安全 - 效用前沿。
AI 中文摘要
微调即服务(FTaaS)平台允许用户在定制任务上训练大语言模型(LLMs),但此流程可能会削弱模型的安全对齐。实际中,服务提供商需要在不重新运行完全对齐或破坏定制任务所获效用的情况下恢复模型安全。现有工作采用模型参数合并,在微调模型参数上添加安全补丁以消除不安全倾向。然而,基于合并的范式受任务 - 安全更新纠缠限制,合并强度难校准。我们将基于合并方法的重点从设计在线合并算子转向离线补丁学习,提出TRACE框架,通过模拟有害调整轨迹生成渐进损坏状态,优化插件补丁在保持效用的同时恢复安全。在六个基准测试和两个模型上,TRACE始终占据安全 - 效用前沿,在所有设置下达到近100%安全,同时保持与未防护微调模型相当的效用。
英文摘要
Fine-Tuning-as-a-Service (FTaaS) platforms let users perform supervised fine-tuning (SFT) on customized data, but this pipeline can erode model safety alignment. To recover safety without re-running full alignment, existing realignment methods focus on calibrating the integration of safety patches into fine-tuned models. These methods exhibit a persistent safety-utility trade-off: weak repair leaves harmful behavior intact, while stronger repair increasingly damages the benign task. This paper shifts the focus from online calibration to offline patch learning and aims to learn a safety patch that restores safety while preserving task-specific capabilities. To this end, we propose TRACE, which simulates harmful SFT trajectories to produce progressively corrupted model states, and optimizes a safety patch simultaneously across these states. TRACE trains the safety patch during the offline stage, and reuses it across all user checkpoints without per-user calibration. We evaluate two representative models using three harmful SFT datasets, together with three utility benchmarks. Across six benchmarks and two models, TRACE consistently dominates the safety-utility frontier. TRACE improves the safety rate by up to 77 percentage points over the best baselines, while maintaining comparable utility to the undefended fine-tuned model.