发表机构
Google Cloud AI Research; Google; Michigan State University(谷歌云人工智能研究; 谷歌; 密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体面临间接提示注入风险,提出ASPIRE红队测试引擎,通过行为图与探索/利用专家发现并验证漏洞,显著扩展覆盖范围并保持高攻击成功率。
AI 中文摘要
大型语言模型智能体检索不可信内容并通过工具执行操作,从而产生间接提示注入风险,可能导致未经授权的操作或持久性状态更改。现有的自动化红队测试主要针对预定义场景优化载荷,未能探索智能体行为空间中潜在的漏洞。我们提出了ASPIRE,一个用于开放式、行为级漏洞发现的智能体安全与提示注入红队测试引擎。ASPIRE维护一个不断演进的智能体安全行为图,并使用互补的探索和利用专家来发现、验证和泛化以后果为中心的测试。轨迹证据更新该图并诊断部分或失败的尝试,而跨运行记忆则传递有效的红队策略。在多个基准上的实验表明,ASPIRE在后果、注入方法、环境和行为路径上的覆盖范围大幅扩展,同时保持了较高的攻击成功率。
英文摘要
LLM agents retrieve untrusted content and act through tools, creating indirect prompt-injection risks that can cause unauthorized actions or persistent state changes. Existing automated red-teaming largely optimizes payloads for pre-specified scenarios, leaving latent vulnerabilities across the agent's behavior space unexplored. We present ASPIRE, an Agentic Safety & Prompt Injection Red-teaming Engine for open-ended, behavior-level vulnerability discovery. ASPIRE maintains an evolving Agent Security Behavior Graph and uses complementary Explore and Exploit experts to discover, verify, and generalize consequence-centric tests. Trajectory evidence updates the graph and diagnoses partial or failed attempts, while cross-run memory transfers useful red-team strategies. Experiments on various benchmarks show that ASPIRE substantially expands coverage across consequences, injection methods, environments, and behavior paths while maintaining strong attack success.