arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体攻击智能体:用于生产智能体红队测试的自动研究

Agent Hacks Agents: Autoresearch Discovers Vulnerabilities in Production Agents

Xutao Mao, Rui Qian, Xiang Zheng, Cong Wang

arXiv 2607.11698首次发表:更新:

发表机构

City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究生产型语言模型智能体的自动化红队测试,提出AHA可证伪发现循环,通过该循环构建漏洞概念图。在Claude Code和Codex上测试,发现跨模型和智能体的可重复使用漏洞核心,生成的VCG为安全团队提供可审计工件,提升测试性能。

AI 中文摘要

生产型语言模型智能体(如Claude Code和Codex)在不受信任的内容、文件、命令和工作区状态上运行,安全失败可直接导致行动。现有红队测试方法主要优化攻击成功率并保留基准、有效载荷等工件,未记录不安全智能体行为背后的促成条件。我们使用一个智能体研究环境来研究生产型语言模型智能体的自动化红队测试,以发现关于另一个智能体的可重复使用的漏洞知识。我们提出了AHA,这是一个可证伪的发现循环,它提出漏洞假设、构建证伪器、实例化有效攻击、在沙盒环境中执行并反思轨迹,将确认的发现纳入漏洞概念图(VCG)。每个概念通过断言、促成条件、证伪器、转移预测和支持证据将面向攻击者的表面与不安全轨迹联系起来。在Claude Code和Codex上针对三种涵盖直接和间接攻击的场景进行测试,发现的概念揭示了跨模型和智能体的可重复使用的漏洞核心。一个冻结的VCG无需进一步搜索,在相同的单次协议下,比最强的冻结发现基线性能高出14.2个百分点,同时可跨场景和攻击渠道转移。生成的VCG为生产安全团队提供了一个可审计的工件,用于检查漏洞、验证补丁并积累可重复使用的安全知识。

英文摘要

Production LLM agents such as Claude Code and Codex can modify files and execute commands, so safety failures become real destructive actions. Automatic red-teaming lets safety teams test beyond static suites at the pace of deployment updates. Existing methods retain successful attacks but not why each attack succeeded, so after a failed reuse testers cannot tell whether the weakness is gone or the attack no longer fits. We instead retain vulnerability concepts, each stating why an attack succeeds, the condition that enables it, and the evidence that would refute it. Agent Hacks Agents (AHA) discovers these concepts with a Karpathy-style autoresearch loop that tests falsifiable hypotheses on agent trajectories. Repeatedly confirmed concepts enter a vulnerability concept graph that links related weaknesses. Across 18 settings of three scenarios, three victim models, and two agents, the concepts divide into eight families. A shared core recurs across agents, while diverse families appear only in specific settings. On held-out instances the concepts reach a 47.0% attack success rate (ASR) against 32.8% for the strongest baseline, which uses more discovery queries. Beyond the discovery setting, the concepts support two kinds of reuse. Generalization carries concepts to new victims, scenarios, and harnesses, and coordination joins graph-linked concepts into stronger attacks on two of three scenarios. For defense, patching each concept's enabling condition lowers AgentHazard ASR by 41.11 points. The concepts explain why production agents fail and where to repair the agents. Our code is in https://github.com/henrymao2004/Auto-research-red-teaming.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑