arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03509cs.CR

SkillJack:自进化智能体中的持续性技能后门

SkillJack: Persistent Skill Backdoors in Self-Evolving Agents

Zonghao Ying, Xiangfan Wu, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo

AI总结:

该研究提出SkillJack攻击,利用自进化智能体的经验转技能流水线植入持续性恶意技能,在SkillX和Anything2Skill上验证其低检测率、高成功率,揭示技能演化为新攻击面。

AI中文摘要:

自进化智能体正日益将交互历史转化为可复用的、可在单个任务之外持续存在的技能。现有研究关注内存和检索污染,但这类攻击仅在中毒记录被作为上下文检索时才会影响智能体。我们揭示了一种新的、更根本的风险:中毒的经验可被智能体自身转化为持久的行为产物。我们提出SkillJack,这是首个利用自进化智能体的经验到技能流水线的攻击。SkillJack不直接操纵运行时上下文,而是劫持智能体自身的学习过程,将恶意行为植入其可复用的技能库中。我们确定了这种转化的三个关键特性:“清理洗白”,即恶意意图在技能提取过程中被掩盖;“跨层提升”,即短暂的经验成为持久的能力;“持久性隔离”,即攻击在其原始源记录被删除后仍然存在。我们在两个代表性系统SkillX和Anything2Skill上评估SkillJack,使用包含四个策略风险类别的150条轨迹的共享数据集。结果显示,技能提取大幅降低了攻击的可检测性:在SkillX中,安全检测从中毒轨迹的98.5%降至提取技能的11.4%,而Anything2Skill呈现类似效果。同时,植入的技能仍然有效,在两个系统上分别达到56.2%和89.2%的攻击成功率。此外,80.0%的技能介导攻击在删除原始中毒记录后仍然存在,且部分技能会在良性查询上意外激活。我们的发现揭示技能演化为新的攻击面,并推动了对感知来源的技能生命周期保护。我们的代码可在该https URL获取。

英文摘要:

Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only affect agents when poisoned records are retrieved as context. We uncover a new and more fundamental risk: poisoned experiences can be transformed by the agent itself into durable behavioral artifacts. We present \textbf{SkillJack}, the first attack that exploits the experience-to-skill pipeline of self-evolving agents. Instead of directly manipulating runtime context, SkillJack hijacks the agent's own learning process to implant malicious behaviors into its reusable skill repertoire. We identify three key properties of this transformation: \emph{sanitization whitewashing}, where malicious intent is obscured during skill extraction; \emph{cross-layer promotion}, where transient experiences become persistent capabilities; and \emph{persistence isolation}, where the attack survives removal of its original source records. We evaluate SkillJack on two representative systems, SkillX and Anything2Skill, using a shared dataset of 150 trajectories across four policy-risk categories. Results show that skill extraction substantially reduces attack detectability: in SkillX, safety detection drops from 98.5\% for poisoned trajectories to 11.4\% for extracted skills, while Anything2Skill shows a similar effect. Meanwhile, the implanted skills remain effective, achieving attack success rates of 56.2\% and 89.2\% on the two systems, respectively. Furthermore, 80.0\% of skill-mediated attacks persist after deleting the original poisoned records, and some skills unintentionally activate on benign queries. Our findings reveal skill evolution as a new attack surface and motivate provenance-aware skill lifecycle protection. Our code is available at https://github.com/Tencent/AI-Infra-Guard/research/skilljack.

↑