大语言模型智能体中的后门净化动力学
Backdoor Decontamination Dynamics in LLM Agents
浏览论文内容
中文总结 AI 辅助
本研究针对大语言模型智能体的后门问题,提出框架研究净化动力学,发现防御性投毒加遗忘可高效擦除多数后门,中间层仍留触发器感知痕迹。
中文摘要 AI 辅助
开源权重大语言模型(LLM)智能体易在微调过程中被植入后门,若测试阶段从未满足触发条件则可能无法被检测到。假设防御者不知道现有触发器,无法直接对其进行遗忘。一种净化策略是植入已知后门(防御性投毒)后再进行遗忘,希望作为副作用移除原本未知的后门,但该过程结果不确定:原后门可能持续存在、被擦除或转向等。我们引入了一个用于研究工具调用智能体中此类动力学的框架,在AgentDyn上通过系统实验将触发器、响应、教师模型和微调方法解耦。在115项实验中,仅防御性投毒就能擦除约56%的原后门;随后的净化几乎将所有剩余后门擦除,证实触发器识别与恶意执行在行为上可分离。有趣的是,我们的实验发现,当防御性投毒后采用遗忘进行净化时,与防御后门属于同一大类的不同触发器的恶意后门从未持续存在。共同植入最多四个后门会增加抗性(约36%被擦除),但净化单个已知共同驻留后门会附带清除52/60个共同驻留后门(87%)。使用J-lens可视化净化后的模型内部,我们确认尽管净化恢复了良性LLM响应,但中间层仍保留了原触发器感知的痕迹。
英文摘要
Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defenders do not know the existing trigger, they cannot unlearn it directly. One decontamination strategy is to install a known backdoor (defensive poisoning) then to unlearn it, hoping that the original unknown backdoor is removed as a side effect. However, this procedure has uncertain outcomes: the original backdoor may persist or be erased or rerouted, among other possibilities. We introduce a framework for studying these dynamics in tool-calling agents, decoupling trigger, response, teacher, and fine-tuning method across systematic experiments on AgentDyn. Across 115 experiments, defensive poisoning alone erases around 56% of original backdoors; subsequent decontamination then drives almost all survivors to erasure, confirming that trigger recognition and malicious execution are behaviorally dissociable. Interestingly, our experiments find that malicious backdoors never persist when using different triggers of the same general type as the defensive backdoor when followed by decontamination via unlearning. Co-installing up to four backdoors increases resistance (around 36% erased), yet decontaminating a single known co-resident backdoor collaterally clears 52/60 co-residents (87%). Upon visualizing postdecontamination model internals using J-lens, we confirm that although the decontamination restores benign LLM responses, traces of original trigger awareness persist at intermediate layers.
发表机构
- ServiceNow Research(ServiceNow研究院)
- Mila -- Quebec AI Institute(米拉-魁北克人工智能研究所)
- McGill University(麦吉尔大学)
- Canada CIFAR AI Chair(加拿大CIFAR人工智能讲席)
机构由 AI 辅助整理,请以论文原文为准。