arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EVOMAL:自进化编码智能体中的自中毒

EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

Xiaodong Wu, Yu Shi, Qi Li, Zhimin Zhao, Xiangman Li, Bram Adams, Ahmed E. Hassan, Jianbing Ni

arXiv 2608.25776首次发表:更新:

发表机构

Queen’s University(女王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现自进化编码智能体存在自中毒漏洞,提出EvoMal攻击可实现自传播恶意技能,而counter-prompt防御能有效降低攻击成功率且不影响任务完成。

AI 中文摘要

自进化大语言模型(LLM)编码智能体通过模仿从共享技能库中检索到的技能来编写自身工具。我们在该循环中发现了一个漏洞:在创作过程中,检索到的恶意技能可成为保留有效载荷的新技能模板,我们将此称为自中毒:智能体创作、存储并运行由此产生的恶意技能。我们通过EvoMal(一种攻击方式)利用该漏洞,该攻击通过将可互换有效载荷包装在banner(一组看似良性的结构元素,可诱导模仿智能体重现所附代码)中放大自中毒效果。攻击者在不调用恶意技能的情况下将其植入技能库,随后智能体创作并执行携带有害代码的新技能,每个创作的副本可重新进入技能库并被再次模仿,形成在植入技能被移除后仍持续存在的自传播蠕虫。我们定义智能体自中毒率(ASPR)为将新创作的恶意技能添加到库中的任务比例。在153个与工具相关的SWE-bench Verified任务上,针对六种模型的ASPR范围为20.3%至41.8%,中毒库中的恶意技能数量是植入量的4.9至9.0倍。该漏洞在无banner时也会出现:DeepSeek-V4-Pro仅通过有效载荷即可达到11.1%的ASPR;针对某一任务系列定制植入技能描述可将ASPR提升至86.7%;在植入技能被移除后,Qwen3在第5轮仍保持68%的ASPR,因为智能体创作的副本依然存在,这些副本可规避现有防御措施(现有防御措施聚焦于攻击者提交的名称、代码和签名)。我们提出counter-prompt(一种防御方法),其可抑制banner式复制,将EvoMal的ASPR降至最高6.7%,且不会显著降低任务完成率。

英文摘要

Self-evolving LLM coding agents write their own tools by imitating retrieved skills from shared skill libraries. We identify a vulnerability in this loop: during authoring, a retrieved malicious skill can become the template for a new skill that preserves the payload. We call this self-poisoning: the agent authors, stores, and runs the resulting malicious skill. We exploit it through EvoMal, an attack that amplifies self-poisoning by wrapping an interchangeable payload in a banner, a set of benign-looking structural elements that induces an imitating agent to reproduce the enclosed code. The attacker plants malicious skills in the library without invoking them. The agent then authors and executes new skills carrying the harmful code. Each authored copy can re-enter the library and be imitated again, forming a self-propagating worm that persists after the planted skills are removed. We define the agent self-poisoning rate (ASPR) as the fraction of tasks that add a newly authored malicious skill to the library. Across six models on 153 tool-relevant SWE-bench Verified tasks, ASPR ranges from 20.3% to 41.8%, and the poisoned libraries hold 4.9 to 9.0 times as many malicious skills as were planted. The vulnerability also appears without a banner: DeepSeek-V4-Pro reaches 11.1% ASPR with the payload alone. Tailoring the planted skill descriptions to one task family raises ASPR to 86.7%. After the planted skills are removed, Qwen3 retains a round-5 ASPR of 68% because agent-authored copies remain. These copies evade existing defenses, which focus on attacker-submitted names, code, and signatures. We propose counter-prompt, a defense that discourages banner-style copying and reduces EvoMal's ASPR to at most 6.7% with no significant task-completion loss.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑