arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27496cs.CR

ROPE:针对间接提示注入的路由来源策略执行

ROPE: Routed Origin Policy Enforcement against Indirect Prompt Injection

Xinhang Ma, Chaowei Xiao, William Yeoh, Ning Zhang, Yevgeniy Vorobeychik

AI总结:

本文提出ROPE防御间接提示注入,通过确定性来源检查实现可证明的安全保证,在低攻击成功率的同时保留高智能体效用,优于现有系统级防御。

AI中文摘要:

间接提示注入(IPI)会将指令植入使用工具的大语言模型(LLM)智能体所读取的内容中,引导智能体执行有害的工具调用。最强的防御措施是系统级的,利用诸如任务条件工具筛选以防止执行恶意工具,以及信息流控制以避免使用不可信参数执行工具。然而,随着智能体能力增强,用户将更多任务委托给自动化处理,导致工具执行序列和参数值越来越多地在运行时确定,若仅基于用户查询进行筛选,会造成显著的效用损失,无法可靠完成筛选。本文提出ROPE(Routed Origin Policy Enforcement,路由来源策略执行),其基于一种结构性信任概念:一个值只有在可不可伪造地追溯到用户、用户明确指定的来源或用户自己的权威记录时,才能到达状态变更工具。执行过程是对经审计的敏感工具参数集合进行确定性来源检查,且仅对语言模型的依赖限于用户的可信请求,该请求在攻击者的控制范围之外。该方法具有两个可证明的保证:1)在轨迹的每一步,任何仅起源于攻击者可写内容的值都不会到达受来源保护的参数;2)注入内容的任何改写都不会改变准入决策。我们在开放式智能体套件上对四种智能体模型进行评估,结果显示ROPE将攻击成功率控制在1.6%至2.6%之间,同时保留了未防御状态下82%至100%的干净效用,在效用方面显著超越最先进的系统级防御措施,同时达到相当或更好的安全性。此外,我们表明针对ROPE优化注入基本无效,而能击败先前系统级防御的长程攻击成功率为0。我们的代码和日志可在该https URL获取。

英文摘要:

Indirect prompt injection (IPI) plants instructions in the content a tool-using LLM agent reads, steering the agent into harmful tool calls. The strongest defenses are system-level, leveraging techniques such as task-conditional tool screening to prevent execution of malicious tools, and information-flow control to avoid tool execution with untrusted parameters. However, as agents grow more capable, users delegate more to automation. Consequently, tool execution sequences and parameter values are increasingly determined at runtime and cannot be reliably screened from solely user's query without significant utility loss. We present ROPE (Routed Origin Policy Enforcement), which is anchored in a structural notion of trust: a value may reach a state-changing tool only if it traces unforgeably to the user, a source the user explicitly named, or the user's own authoritative records. Enforcement is then a deterministic origin check over an audited set of sensitive tool parameters, and the only reliance on a language model involves solely the trusted user request, out of the attacker's reach. Our approach admits two provable guarantees: 1) at every step of a trajectory, no value whose only origin is attacker-writable content reaches an origin-guarded parameter, and 2) no rewording of an injection changes an admission decision. We evaluate across four agent models on open-ended agent suites, ROPE holds attack success rate to 1.6--2.6\% while retaining 82--100\% of undefended clean utility, significantly exceeding state-of-the-art system-level defenses in utility while attaining comparable or better security. Further, we show that optimizing the injection against ROPE is largely ineffective, while long-horizon attacks that defeat prior system-level defenses achieve zero success rate. Our code and logs are available at https://github.com/xhOwenMa/ROPE .

↑