arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ASCENT:面向安全性与效用协同提升的一阶最优微调与重校准方法

ASCENT: First-Order Optimal Fine-Tuning with Recalibration for Safety--Utility Co-Enhancement

Weiwei Qi, Chongyu Wang, Tianhang Zheng, Zefeng Wu, Zhilin Guo, Xiaojun Jia, Zhongjie Ba, Kui Ren

arXiv 2610.08061首次发表:更新:

发表机构

The State Key Laboratory of Blockchain and Data Security, Zhejiang University; Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security; Nanyang Technological University(浙江大学区块链与数据安全国家重点实验室; 杭州高新技术开发区(滨江)区块链与数据安全研究院; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ASCENT提出一阶最优安全感知校准与任务更新交替的微调框架,理论刻画最优安全子空间,在多个LLM上实现效用提升20.3%、攻击成功率降低35.5%,达成安全与效用协同增强。

AI 中文摘要

监督式微调可以显著提升大型语言模型(LLMs)的下游效用,但可能损害其安全性。现有的安全性保持方法使用与安全相关的参数或子空间来约束下游更新,但主要关注安全性保持而非安全性与效用的联合提升,缺乏对最优安全性相关子空间和安全性保持任务更新的理论刻画,并且通常依赖静态安全性子空间,该子空间在微调过程中可能变得过时。为解决这些局限性,我们提出了ASCENT,一种通过一阶最优安全性感知周期性校准和任务优化来实现安全性与效用协同增强的下游微调框架。我们将安全性建模为LLM参数$S(\ heta)$的函数,并使用其一级近似来刻画参数更新下的安全性变化。在固定秩和Frobenius范数预算下,我们证明了由安全性函数梯度的前$r$个奇异分量构建的更新能够最大化估计的安全性变化,并将其用于周期性校准以保持和提升安全性。我们进一步推导出一个独特的安全性保持任务更新,该更新在保持接近原始任务更新的同时,惩罚对估计安全性变化的负面影响。ASCENT交替执行这些最优任务更新和校准更新,以联合增强安全性和效用。在多个LLM家族和下游任务上的实验表明,ASCENT将下游效用提升高达20.3%,并将攻击成功率降低高达35.5%,在所有评估设置中实现了最先进的安全性和效用。我们的代码可在https URL获取。

英文摘要

Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but may compromise their safety. Existing safety-preserving methods constrain downstream updates using safety-related parameters or subspaces, but mainly focus on safety preservation rather than joint safety and utility enhancement, lack a theoretical characterization of the optimal safety-related subspace and safety-preserving task update, and typically rely on a static safety subspace that may become outdated during fine-tuning. To address these limitations, we propose ASCENT, a downstream fine-tuning framework for safety--utility co-enhancement through first-order optimal safety-aware periodic calibration and task optimization. We model safety as a function of LLM parameters $S(θ)$ and use its first-order approximation to characterize safety changes under parameter updates. Under a fixed rank and Frobenius-norm budget, we prove that the update constructed from the top-$r$ singular components of the safety-function gradient maximizes the estimated safety change, and use it for periodic calibration to preserve and improve safety. We further derive a unique safety-preserving task update that stays close to the original task update while penalizing negative effects on the estimated safety change. ASCENT alternates these optimal task and calibration updates to jointly enhance safety and utility. Experiments across multiple LLM families and downstream tasks show that ASCENT improves downstream utility by up to 20.3\% and reduces attack success rate by up to 35.5\%, achieving state-of-the-art safety and utility across all evaluated settings. Our code is available at https://github.com/ZJU-LLM-Safety/ASCENT.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑