发表机构
University of Illinois Urbana-Champaign; MATS Research; University of Oxford, OATML; Measuring AI Progress, Inc.(伊利诺伊大学厄巴纳-香槟分校; MATS 研究; 牛津大学 OATML; AI 进步测量公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示LLM智能体后训练中后门可存活,并提出PersistBD方法增强后门持久性,在Qwen2.5-Coder-7B上将攻击成功率从20%提升至76%,警示供应链风险。
AI 中文摘要
开发者可以通过良性后训练来适配第三方模型,从而构建LLM智能体。我们研究了一种供应链威胁,其中攻击者提供带有后门的模型:当特定输入模式出现时,该模型会产生恶意输出的隐藏行为。我们聚焦于软件工程智能体,探究此类后门是否能在开发者的监督微调(SFT)及随后的任务级强化学习(RL)中存活。我们观察到,良性SFT大幅降低了攻击成功率,但随后的RL通常会保留残余行为,有时甚至会增加攻击成功率。我们对SFT期间后门侵蚀的分析确定了两个可能有利于后门存活的因素:初始后门强度和与良性训练的梯度兼容性。这些因素促使我们提出PersistBD,它在模型发布前对已植入后门的模型进行优化,以提高其在良性后训练过程中的持久性。在Qwen2.5-Coder-7B上,PersistBD将SFT后的攻击成功率从20%提升至74%,并将SFT-RL后的攻击成功率从20%提升至76%,同时保持了相当的良性任务性能。综合来看,我们的结果表明,后门可以在良性后训练中保持活跃,且攻击者可以有意提高其持久性。这凸显了AI开发者的供应链风险,并促使在将第三方模型适配为智能体时,需要更强的技术来检测和缓解继承的后门。我们的代码可在该https URL获取。
英文摘要
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level reinforcement learning (RL). We observe that benign SFT substantially reduces attack success, but subsequent RL often preserves the residual behavior and sometimes even increases attack success. Our analysis of backdoor erosion during SFT identifies two factors that may favor survival: initial backdoor strength and gradient compatibility with benign training. These factors motivate PersistBD, which refines an already-backdoored model before release to improve its persistency through the benign post-training process. On Qwen2.5-Coder-7B, PersistBD raises attack success from 20% to 74% after SFT and from 20% to 76% after SFT-RL, while maintaining comparable benign task performance. Together, our results show that backdoors can remain active through benign post-training and that adversaries can deliberately increase their persistence. This highlights a supply-chain risk for AI developers and motivates stronger techniques for detecting and mitigating inherited backdoors when adapting third-party models into agents. Our code is available at https://github.com/uiuc-kang-lab/PersistBD.
CommentsEMNLP 2026 Findings