发表机构
Key Lab of HCST (PKU), MoE and School of Computer Science , Peking University(北京大学计算机科学与技术学院,教育部复杂系统管理与控制国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对编码智能体易受间接提示注入攻击的问题,提出CapScope框架级授权机制,通过能力隔离限制工具调用,在保持修复成功率的同时显著降低注入执行率。
AI 中文摘要
编码智能体使用系统级工具来读取文件、执行命令和修改源代码。在智能体的沙箱内,这些工具通常携带环境权威:命名一个资源就足以对其执行操作。间接提示注入通过将指令放置在仓库文件或工具输出中,导致智能体执行用户未请求的操作,从而利用这种权威。我们提出CapScope,一种框架级授权机制,在不要求模型识别恶意文本的情况下限制工具使用。在读取仓库内容或工具输出之前,CapScope从受信任输入中推导出任务范围的权威上限。然后,它为每个智能体分配一组独立的类型化能力,存储于模型上下文之外。每次工具调用都会根据发出调用的智能体的能力进行检查。因此,分配给一个子智能体的权限不会自动对其他子智能体可用。注入可能导致智能体请求某个操作,但除非该智能体已具备所需能力,否则该请求将被阻止。我们在Pi编码智能体上实现CapScope,并在一个修复工作流中进行评估,其中编排器将子任务委派给独立的子智能体。评估覆盖五个Python任务、五个注入面、四种授权条件和每单元三次试验(共300次运行)。在环境权威和全局策略基线条件下,注入效果在33-47/75次运行中执行,而CapScope下仅为3/75次。CapScope完成68/75次修复,而基线完成68-72/75次。
英文摘要
Coding agents use system-level tools to read files, execute commands, and modify source code. Within the agent's sandbox, these tools often carry ambient authority: naming a resource is sufficient to act on it. Indirect prompt injection exploits this authority by placing instructions in repository files or tool output that cause the agent to perform actions the user did not request. We propose CapScope, a harness-level authorization mechanism that restricts tool use without requiring the model to identify malicious text. Before repository contents or tool output are read, CapScope derives a task-wide authority ceiling from trusted input. It then assigns each agent a separate set of typed capabilities, stored outside the model's context. Every tool call is checked against the capabilities of the agent that issued it. Permissions assigned to one sub-agent are therefore not automatically available to another. An injection may cause an agent to request an action, but the request is blocked unless that agent already has the required capability. We implement CapScope on the Pi coding agent and evaluate it in a repair workflow where an orchestrator delegates subtasks to separate sub-agents. The evaluation covers five Python tasks, five injection surfaces, four authorization conditions, and three trials per cell (300 runs). The injected effect executes in 33-47/75 runs under the ambient-authority and global-policy baselines, compared with 3/75 under CapScope. CapScope completes 68/75 repairs, while the baselines complete 68-72/75.