arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPA:通过优先计划的信息流控制保护跨查询的持久化大语言模型智能体

SPA: Securing Persistent LLM Agents Across Queries with Plan-First Information-Flow Control

Dylan Girrens, Guangjing Wang

arXiv 2608.27234首次发表:更新:

AI 中文总结

SPA是一种优先计划的架构,通过双格信息流控制保护LLM智能体的规划、执行与跨查询状态复用,在AgentDojo和AgentDojo-MQ上可显著降低攻击成功率,同时揭示了安全与效用的权衡。

AI 中文摘要

大语言模型(LLM)智能体越来越多地在不可信的网页、文档、工具和持久化状态上运行,同时对安全敏感资源拥有控制权。现有防御措施通常仅保护规划过程或单个工具交互,但持久化智能体面临更广泛的威胁:攻击者控制的数据可改变控制流、注入安全敏感的工具参数,或危害后续查询。我们提出SPA,一种优先计划的架构,用于保护规划、执行和跨查询状态复用。SPA为每个查询调用一次规划器,以声明式领域特定语言生成完整的可执行计划,随后应用双格信息流控制,跟踪显式数据流和控制依赖关系的保密性与完整性。为支持持久化且不将不可信负载重新暴露给规划器,SPA将执行结果存储为带标签的工件,仅在后续规划中披露语义元数据。我们在AgentDojo及我们用于测量安全状态复用和延迟攻击的多查询扩展AgentDojo-MQ上评估SPA。在“tool_knowledge”攻击下,采用信息流控制的SPA在AgentDojo上将攻击成功率降至零,在AgentDojo-MQ上降至0.2%。我们的结果表明,优先计划执行结合保留标签的持久化可显著增强持久化LLM智能体,同时揭示了严格完整性执行带来的重要安全-效用权衡。

英文摘要

Large language model (LLM) agents increasingly operate over untrusted webpages, documents, tools, and persistent states while exercising authority over security-sensitive resources. Existing defenses typically protect either planning or individual tool interactions, but persistent agents face a broader threat: attacker-controlled data can alter control flow, enter security-sensitive tool arguments, or compromise later queries. We present SPA, a plan-first architecture that secures planning, execution, and cross-query state reuse. SPA invokes the planner once per query to generate a complete executable plan in a declarative domain-specific language, then applies dual-lattice information-flow control to track confidentiality and integrity across explicit data flows and control dependencies. To support persistence without re-exposing untrusted payloads to the planner, SPA stores execution results as labeled artifacts and reveals only semantic metadata during later planning. We evaluate SPA on AgentDojo and AgentDojo-MQ, which is our multi-query extension for measuring secure state reuse and delayed attacks. Under the 'tool_knowledge' attack, SPA with information-flow control reduces attack success to zero on AgentDojo and 0.2% on AgentDojo-MQ. Our results show that plan-first execution combined with label-preserving persistence can substantially strengthen persistent LLM agents, while revealing an important security-utility tradeoff introduced by strict integrity enforcement.

Comments24 pages, 17 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑