发表机构
Stony Brook University; Amazon(石溪大学; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文揭示在线策略蒸馏用于安全时,带后门的教师模型能以极低投毒率向学生模型传播恶意行为,并发现训练轮数和top-k KL会放大风险,提出惰性防御缓解。\n
AI 中文摘要
在线策略蒸馏(OPD)作为一种将教师模型能力迁移至学生模型的有效方式,日益受到关注。近期研究进一步探索将OPD作为提升大型语言模型安全性的工具,并取得了有前景的结果。然而,这些方法通常假设教师模型和训练数据是可信的。本文揭示了一个被忽视的OPD安全威胁:一个经过安全对齐但带有后门的教师模型,能够将其隐藏的恶意行为传播给最初干净的学生模型。在我们的威胁模型下,仅3%的投毒率即可使蒸馏后的学生模型上的攻击成功率(ASR)高达70%。我们进一步识别出两种可能放大此风险的训练选择。首先,增加训练轮数可在低投毒率下导致高ASR。仅用10个投毒样本,经过16轮训练后ASR即达到67%。其次,常用的top-k KL散度可加速后门迁移,在大多数设置下使触发条件的有害行为比采样token KL更早出现。除这些发现外,我们探索了一种简单的缓解措施——惰性防御(Lazy Defense),它通过裁剪KL奖励使学生更新不那么激进,限制激进更新并减缓后门学习。实验表明,惰性防御在低投毒率设置下能延迟后门迁移。综合来看,我们的发现揭示了OPD可能传播后门,凸显了解决OPD安全风险的必要性。
英文摘要
On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Recent studies further explore OPD as a tool for improving large language model safety with promising results. However, these approaches typically assume that the teacher and training data are trustworthy. In this paper, we uncover an overlooked threat to OPD for safety: a safety-aligned but backdoored teacher can propagate its hidden malicious behavior to an initially clean student. Under our threat model, a poisoning rate as low as 3% results in an attack success rate (ASR) of up to 70% on the distilled student. We further identify two training choices that can amplify this risk. First, increasing the number of training epochs can lead to high ASR even at low poisoning rates. With only 10 poisoned samples, ASR reaches 67% after 16 epochs. Second, the commonly used top-k KL can accelerate backdoor transfer, causing trigger-conditioned harmful behavior to emerge earlier than sampled-token KL in most settings. Alongside these findings, we explore a simple mitigation, Lazy Defense, which clips KL rewards to make student updates less aggressive, limiting aggressive updates and slowing backdoor learning. Experiments show that Lazy Defense delays backdoor transfer in low poisoning rate settings. Together, our findings reveal that OPD can propagate backdoors, highlighting the need to address the safety risks of OPD.