arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当经验成为指令:自进化智能体技能系统中的轨迹中毒攻击

When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems

Jialuo Chen, Lingqi Jiang, Xinhao Deng, Xiaohu Du, Jianan Ma, Yunhao Feng, Yuqi Qing, Zhihao Yuan, Linkang Du, Jingyi Wang

arXiv 2608.05563首次发表:更新:

发表机构

Ant Group; Zhejiang University; Tsinghua University; Hangzhou Dianzi University; The Chinese University of Hong Kong, Shenzhen; Xi’an Jiaotong University(蚂蚁集团; 浙江大学; 清华大学; 杭州电子科技大学; 香港中文大学(深圳); 西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PoisonedEvolution轨迹中毒攻击,针对自进化智能体技能系统的经验转指令过程,在SkillClaw等平台实现高成功率,证明了证据转化是这类系统的安全边界。

AI 中文摘要

自进化技能(SES)系统将智能体轨迹提炼为持久技能,使得不可信的经验能够转化为可信的指令。本文提出PoisonedEvolution,一种针对该转化过程的轨迹中毒攻击方法。该攻击是技能可见的黑盒攻击者,可检查目标技能并提供有限证据,但无法观察私有池、进化逻辑或编辑技能库。人工制品中毒需要满足包含性、进化归因和实现三个条件,其中归因是关键瓶颈:目标行为必须在转化前呈现出因果有用性、可重复性和可泛化性。我们使用惰性金丝雀规范评估了四类代表性安全效果。在SkillClaw中的六个主流LLM进化器上,当攻击者支持率为10%时,PoisonedEvolution在600次试验中将目标行为嵌入了546次,达到91.0%的 SER;在结构不同的Trace2Skill流水线中,以相同比例攻击时,其在600次试验中嵌入了369次,达到61.5%的 SER,证明了该攻击在不同进化架构间的可迁移性。在一项代表性对照研究中,30条记录的批次中仅需3条一致的攻击者记录即可成功,而单条记录的效果则弱得多。消融实验表明,可重复支持、因果框架和领域对齐编码是攻击成功的主要决定因素。这些发现揭示了证据转化是自进化智能体的一个安全边界。

英文摘要

Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion process. Our skill-visible black-box attacker can inspect a target skill and contribute bounded evidence, but cannot observe private pools or evolution logic or edit the skill bank. Artifact poisoning requires Inclusion, Evolution Attribution, and Realization. Attribution is the distinctive bottleneck: the target behavior must appear causally useful, recurrent, and generalizable before promotion. We evaluate four representative security-effect families using inert canary specifications. At 10% attacker support, across six mainstream LLM evolvers in SkillClaw, PoisonedEvolution embeds target behaviors in 546/600 trials (91.0% SER). On the structurally different Trace2Skill pipeline at the same ratio, it embeds target behaviors in 369/600 trials (61.5% SER), demonstrating transfer across evolution architectures. In a representative controlled study, three consistent attacker records suffice in a 30-record batch, whereas a single record is much weaker. Ablations identify recurring support, causal framing, and domain-aligned encoding as the main determinants of success. These findings expose evidence promotion as a security boundary for self-evolving agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑