Pincer:使用数字孪生实现智能体的资源授权
Pincer: Resource Authorization for Agents using a Digital Twin
浏览论文内容
中文总结 AI 辅助
针对编码智能体权限管理,Pincer利用数字孪生自动学习用户最小权限策略,在资源层防御,兼顾安全与实用性,优于现有基线。
中文摘要 AI 辅助
编码智能体已变得越来越长周期、自主化,依赖通用shell,并维护自身的持久性记忆以实现自我改进。虽然这些能力使智能体变得强大,但也使它们更难防御外部对手。限制这种架构的防御措施——类型化工具、信息流控制或策略预测引擎——会牺牲太多功能而无法被采用。当前部署的智能体(例如Claude、Codex)主要依赖用户介导和自动模式沙箱作为其主要防御手段。在用户介导的沙箱中,用户维护的策略会随时间衰减,重复的权限请求导致用户疲劳,而自动模式的工具调用分类器不学习用户特定策略,且不旨在防御对抗性设置。Pincer是一种在资源层运行的新防御机制,与工具调用层的现有防御(如自动模式)协同工作。Pincer的核心是一个数字孪生,一个隔离上下文模型,自动学习并执行动态的用户特定最小权限策略。数字孪生持续学习用户偏好,使其能够作为用户代理处理智能体的权限请求。为了模拟学习阶段,我们提出了一个新的以用户为中心的数据集,包含遵循多日用户-智能体交互记录的示例。我们的评估表明,与多个基线(包括LLM评判器的变体和Conseca(HotOS '25)的改编)相比,Pincer在安全性和实用性方面均表现强劲。我们重点介绍了Pincer的设计在与其他所有基线相比带来显著安全改进的攻击类型,同时在其他类型的攻击中也优于基线。
英文摘要
Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have also made them harder to defend against external adversaries. Defenses that restrict this architecture --- typed tools, information-flow control, or policy prediction engines --- give up too much functionality to be adopted. Agents deployed today (e.g. Claude, Codex) rely on a combination of user-mediated and automode sandboxing as their primary defense. In user-mediated sandboxing, user-maintained policies decay over time and repeated permission requests cause user fatigue, while auto mode's tool-call classifiers learn no user-specific policy and are not meant to defend against adversarial setups. Pincer is a new defense that operates at the resource layer and works alongside existing defenses at the tool-call layer like the auto mode. At the core of Pincer lies a digital twin, an isolated-context model that automatically learns and enforces dynamic user-specific least-privilege policies. The digital twin keeps continually learning the user's preferences allowing it to act as the user's proxy for the agent's permission requests. To emulate the learning phase, we propose a new usercentric dataset with examples following a multi-day transcript of user-agent interaction. Our evaluation shows that Pincer performs strongly on both security and utility in comparison to several baselines which includes variants of LLM judges and adaptations of Conseca (HotOS '25). We highlight attack types where Pincer's design leads to a significant security improvement compared to all other baselines, while outperforming the baselines even for other types of attacks.
发表机构
- UC Berkeley(加州大学伯克利分校)
- Foothill High School(山麓高中)
机构由 AI 辅助整理,请以论文原文为准。