arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对信任的信任之反思,再探:用投毒基准污染自我修改的AI编码智能体

Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

Franziska Roesner, Tadayoshi Kohno

arXiv 2609.17817首次发表:更新:

发表机构

University of Washington; Georgetown University(华盛顿大学; 乔治敦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将汤普森编译器后门攻击推广至自我修改AI编码智能体,证明投毒基准可诱导其演化出编写漏洞代码的指令,且污染常持续存在,需增强此类智能体的韧性。

AI 中文摘要

汤普森的《对信任的信任之反思》表明,编译器可能被投毒以重新植入其自身的后门,以至于即使重新编译干净的源代码也会重现特洛伊木马。如今,大量编码工作由AI编码智能体完成——而且这些智能体越来越多地生成自身的新版本。我们重新审视当“编译器”是一个自我修改的编码智能体时汤普森的攻击。对手能否向智能体的自我评估和自我改进过程提供投毒基准,以诱导未来版本的智能体在干净的、保留的任务上编写易受攻击的代码?我们针对三个最近提出的自我修改编码智能体实例化了这种攻击:达尔文·哥德尔机(带有我们的实验性修改)、自我改进编码智能体以及超智能体(两者均基本未修改)。我们展示了成功的概念验证:例如,对于由Sonnet 4.5驱动的超智能体,我们的投毒基准导致智能体自我演化出在中性URL获取任务上禁用HTTPS证书验证的指令。从我们的实验中,我们提炼出漏洞、基准、模型和智能体脚手架中足以实现基准投毒攻击的属性。此外,我们表明,即使投毒的智能体随后针对干净基准进行演化,污染也常常持续存在。我们讨论了防御方向,并主张自我修改的编码智能体必须设计得对此类攻击更具韧性。

英文摘要

Thompson's "Reflections on Trusting Trust" showed that a compiler can be poisoned to reinsert its own backdoor, so that even recompiling clean source reproduces the Trojan. Today, substantial coding work is done by AI coding agents -- and increasingly, those agents generate new versions of themselves. We reconsider Thompson's attack when the "compiler" is a self-modifying coding agent. Can an adversary supply poisoned benchmarks to the agent's self-evaluation and self-improvement process to induce future versions of the agent to write vulnerable code on clean, held-out tasks? We instantiate this attack against three recently proposed self-modifying coding agents: the Darwin Gödel Machine (with our experimental modifications), the Self-Improving Coding Agent, and Hyperagents (both substantively unmodified). We demonstrate successful proofs-of-concept: for example, with Hyperagents powered by Sonnet 4.5, our poisoned benchmark leads the agent to self-evolve instructions that disable HTTPS certificate validation on neutral URL-fetching tasks. From our experiments, we distill properties of the vulnerability, benchmark, model, and agent scaffolding that are sufficient to enable a benchmark poisoning attack. Moreover, we show that contamination often persists even when a poisoned agent is subsequently evolved against clean benchmarks. We discuss defensive directions and argue that self-modifying coding agents must be designed to be more resilient to such attacks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑