arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EOPSA:高效在线策略自蒸馏安全对齐

EOPSA: Efficient On-Policy Self-Distilled Safety Alignment

Qirui Liu, Yichen Sun, Yan Wang, Yu Mi, Wei Cao, Yue Shen, Zhixuan Chu, Kui Ren

arXiv 2609.34519首次发表:更新:

发表机构

Zhejiang University; Ant Group(浙江大学; 蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线策略自蒸馏安全对齐中监督崩溃与梯度稀释问题,提出EOPSA,通过自适应rollout调度和选择性蒸馏,削减约50%计算并仅反向传播约2%词元,在安全与推理上超越全词元基线。

AI 中文摘要

在线策略自蒸馏(OPSD)已成为安全对齐的一种有前景的范式,通过从以拒绝导向的特权提示为条件的教师模型中蒸馏,提供密集的、词元级别的监督。然而,我们发现该范式存在严重影响训练效率和通用推理能力的关键低效问题。具体而言,我们诊断出两个根本瓶颈:(1)在扩展的 rollout 过程中监督崩溃,即随着学生生成前缀的延长,教师的纠正效能急剧下降,从而向后期词元注入噪声梯度;(2)由风格偏移引起的梯度稀释,即蒸馏目标被特权提示所引发的与安全无关的风格差异所主导,冲淡了真正的安全信号并损害了基础推理能力。为解决这些问题,我们提出了高效在线策略自蒸馏安全对齐(EOPSA),该方法将计算和梯度预算集中于可靠监督的安全关键词元上。EOPSA 包含两个协调机制:(i)自适应 rollout 调度,该机制由新颖的教师救援率(TRR)指标动态约束生成范围,以严格在可靠监督区间内运行;(ii)选择性蒸馏,该机制过滤掉安全中性词元,将梯度更新限制在安全关键转换上。对高达 32B 参数的推理模型的广泛评估表明,EOPSA 将 rollout 计算削减约 50%,并且仅对约 2% 的词元进行反向传播,在安全合规性和推理保持方面均持续优于全词元蒸馏基线。

英文摘要

On-Policy Self-Distillation (OPSD) has emerged as a promising paradigm for safety alignment, delivering dense, token-level supervision by distilling from a teacher conditioned on refusal-oriented privileged prompts. However, we reveal that this paradigm suffers from critical inefficiencies that degrade both training efficiency and general reasoning capabilities. Specifically, we diagnose two fundamental bottlenecks: (1) supervisory collapse over extended rollouts, where the teacher's corrective efficacy degrades precipitously as the student's generation prefix lengthens, injecting noisy gradients into late-stage tokens; and (2) gradient dilution from stylistic shifts, where the distillation objective is dominated by safety-irrelevant stylistic discrepancies induced by privileged prompting, washing out genuine safety signals and impairing base reasoning. To resolve these issues, we propose Efficient On-Policy Self-Distilled Safety Alignment (EOPSA), which concentrates computational and gradient budgets exclusively on reliably supervised, safety-critical tokens. EOPSA incorporates two coordinated mechanisms: (i) Adaptive Rollout Scheduling, which dynamically bounds the generation horizon guided by a novel Teacher Rescue Rate (TRR) metric to operate strictly within reliable supervision regimes; and (ii) Selective Distillation, which filters out safety-neutral tokens to restrict gradient updates exclusively to safety-pivotal transitions. Extensive evaluations across reasoning models up to 32B parameters demonstrate that EOPSA slashes rollout computation by $\sim$50% and backpropagates through merely $\sim$2% of tokens, consistently outperforming full-token distillation baselines in both safety compliance and reasoning retention.

Comments32 pages, 9 figures. Code and models are available at the project repositories

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑