发表机构
Jilin University; The Hong Kong Polytechnic University(吉林大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SkillPoison框架,通过构建并泛化成功经验而非注入恶意内容,渐进式投毒LLM智能体技能,在三个基准上实现95.71%攻击成功率且经验保持任务正确。
AI 中文摘要
自我改进的LLM智能体越来越多地将成功经验提炼为持久、可复用的技能。现有的技能攻击方法通过向单个经验或提取的技能中注入恶意触发器、行为或虚假事实来破坏这一学习流程。然而,此类攻击容易被检测,且注入的恶意行为往往无法累积为持久技能。在本文中,我们证明即使来自经过验证的成功经验,技能投毒也可能发生,而无需使任何单个轨迹具有恶意性。基于这一见解,我们提出了SkillPoison,一种通过成功经验渐进式投毒技能的新框架。SkillPoison首先构建一组强化目标行为的成功经验,然后移除限制该行为适用场景的上下文条件。SkillPoison并非注入恶意内容,而是塑造技能提取器的泛化方式,使有用行为在支持任务成功的同时,在误用场景下诱导有害行为。在三个基准上的大量实验表明,SkillPoison实现了95.71%的攻击成功率,同时所有注入的经验均保持任务正确性,并通过验证和词法检查。我们的代码、数据和实现细节已公开供社区使用,详见此https链接。
英文摘要
Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks are easily detected, and the injected malicious behaviors often fail to accumulate as persistent skills. In this paper, we show that skill poisoning can arise even from verified successful experiences, without making any individual trajectory malicious. Based on this insight, we propose SkillPoison, a novel framework that progressively poisons skill via successful experiences. SkillPoison first constructs a set of successful experiences that reinforce a target behavior, and then removes the contextual conditions that constrain when the behavior applies. Rather than injecting malicious content, SkillPoison shapes how the skill extractor generalizes, allowing useful behavior to support task success while inducing harmful behavior when they are misapplied. Extensive experiments on three benchmarks show that SkillPoison achieves 95.71% attack success rates, while all injected experiences remain task-correct and pass verification and lexical inspection. Our code, data and implementation details are available for the community at https://github.com/DEEP-JLU/SkillPoison.