发表机构
National University of Singapore; University of North Carolina at Chapel Hill; University of California, Berkeley; University of Chicago; University of California, Santa Barbara(新加坡国立大学; 北卡罗来纳大学教堂山分校; 加州大学伯克利分校; 芝加哥大学; 加州大学圣巴巴拉分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgentXploit提出双角色审计系统,分离仓库级攻击路径发现与运行时利用,在72个漏洞基准上实现59.3%端到端成功率,优于现有方法,揭示智能体安全审计中两类独立挑战。
AI 中文摘要
AI智能体将语言模型与外部数据和工具相结合,这些工具可以修改文件、调用API或执行代码。当对抗性内容改变智能体的工具使用方式,或当周围软件包含路径遍历或命令注入等漏洞时,可能出现安全故障。我们研究授权的白盒部署前审计,其中审计员可以访问目标仓库和受控运行时,但成功的攻击仍必须通过任务定义的攻击者接口进行,并由外部验证器确认。我们提出AgentXploit,一个双角色审计系统,将仓库级攻击路径发现与运行时利用分开。分析器智能体追踪攻击者控制的输入到敏感操作,并记录代码支持的候选攻击路径;利用器智能体将这些路径转化为具体攻击,并使用运行时反馈进行修订。我们还引入了AgentXploit-Bench,包含12个开源AI智能体系统和框架中的72个可复现漏洞。在三次运行中,AgentXploit达到59.3%的端到端成功率,而Codex为38.4%。在令牌预算匹配的比较下,Codex达到46.3%。在AgentDojo上,提供了注入点,利用器智能体达到79.2%的攻击成功率,而AgentVigil为52.7%。这些结果突出了仓库发现和运行时利用作为端到端智能体安全审计中的不同挑战。
英文摘要
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection. We study authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined attacker interface and be confirmed by an external verifier. We present AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation. The Analyzer Agent traces attacker-controlled inputs to sensitive operations and records code-supported candidate attack paths; the Exploiter Agent turns these paths into concrete attacks and revises them using runtime feedback. We also introduce AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks. Across three runs, AgentXploit reaches 59.3% end-to-end success, compared with 38.4% for Codex. Under a token-budget-matched comparison, Codex reaches 46.3%. On AgentDojo, where injection points are provided, the Exploiter Agent reaches 79.2% attack success versus 52.7% for AgentVigil. These results highlight repository discovery and runtime exploitation as distinct challenges in end-to-end agent security auditing.
Comments20 pages, 2 figures