arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

孪生智能体:用于特权分离智能体的上下文残差压缩

Twin Agent: Context Residual Compression for Privilege Separated Agents

Zhanhao Hu, Dennis Jacob, Xiao Huang, Zhaorun Chen, Bo Li, David Wagner

arXiv 2607.19595首次发表:更新:

发表机构

University of California, Berkeley; University of Chicago; University of Illinois, Urbana-Champaign(加州大学伯克利分校; 芝加哥大学; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型智能体易受安全风险问题,提出孪生智能体模式,由探索和安全智能体组成,通过探索智能体传递紧凑提示,减少信息需求,在保持高任务效用的同时防止提示注入攻击,优于其他基线。

AI 中文摘要

大语言模型智能体易受安全风险影响,如来自不可信上下文的提示注入攻击,会操纵下游推理和工具使用。现有设计安全的方法通过分离不可信观测与特权执行及仔细控制信息流来降低风险,但常降低效用且需大量特定任务工程。本文提出孪生智能体,一种受智能体上下文中残差编码启发的通用特权分离设计模式。它由两个近乎对称的智能体组成:探索智能体检查不可信信息,安全智能体执行特权操作。探索智能体以上下安全智能体的当前上下文为条件,仅向安全智能体传达关于下一步行动的紧凑提示。该设计减少了保留任务效用所需的信息,实现了更好的安全-效用权衡,通过测量提示长度变化时效用和攻击成功率的变化进行了实证验证。在软件工程任务和多工具交互任务上进行评估,孪生智能体在防止提示注入攻击的同时保持高任务效用,优于未防护智能体和特权分离基线。

英文摘要

Large language model (LLM) agents are vulnerable to security risks, such as prompt injection attacks from untrusted context that manipulate downstream reasoning and tool use. Existing secure-by-design approaches mitigate this risk by separating untrusted observations from privileged execution and careful control of information flow, but often degrade utility and require extensive task-specific engineering. We thus propose Twin Agent, a general privilege separation design pattern inspired by residual coding in the agent context. Twin Agent consists of two nearly symmetric agents: an Explore Agent that inspects untrusted information and a Safe Agent that executes privileged actions. The Explore Agent is conditioned on the Safe Agent's current context and communicates only compact hints to the Safe Agent about the next action to take. This design reduces the information needed to preserve task utility and thus achieves a better security--utility tradeoff, which we empirically verify by measuring how utility and attack success change as the length of hints varies. We evaluate Twin Agent on long-horizon software engineering tasks with SWE-bench Lite and on heterogeneous multi-tool interaction tasks with AgentDojo and DecodingTrust-Agent. Across both benchmarks, Twin Agent preserves high task utility while preventing prompt injection attacks, outperforming both undefended agents and privilege separation baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑