arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27646cs.AI

如果智能体是天使,就不需要治理:可信工具边界处带外策略执行

If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary

Marc Millstone, Tyler Akidau, Johannes Brüderl, Marat Pekker

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出带外策略执行(OBPE)作为智能体推理外的可信边界,通过原型实验将智能体追踪失败率大幅降低,提升了安全有用的任务完成率。

中文摘要 AI 辅助

为智能体提供人类的凭证,它就会继承该人的权限范围,却没有限制其使用的判断,它可以将所有可访问的记录纳入模型上下文,其中隐藏的指令会引导其下一次调用,并且在智能体超出其职责范围或获取秘密时,每个请求仍保持凭证有效性。提示词是脆弱的护栏:一个易出错的推理器解释任务并执行其限制。我们提出带外策略执行(Out-of-Band Policy Enforcement, OBPE),这是智能体推理之外的可信边界,它对类型化操作和资源进行授权,在后端调用前缩小查询范围,然后过滤响应中的记录和字段或屏蔽值。语义门控可根据参数值或外部状态拒绝或暂停已授权的调用。数据策略所有者设定最大授权,智能体策略只能缩小该范围。我们证明,在规定条件下,策略计划与顺序无关,且智能体策略无法扩大上限。字段移除覆盖一次执行;屏蔽和历史规则要求更少。我们发布了从生产系统简化而来的HTTP代理原型,其一致性测试将类型化的Cedar策略核心与模型关联。针对Jira和ServiceNow的模拟环境,我们的基准测试比较了在四个模型(包括20个自适应红队任务)上,使用和不使用OBPE的提示智能体。追踪失败指受保护数据进入智能体上下文、确切值出现在答案中或完成了禁止的效果。在3621次试验中,失败率从57.6%降至0.2%,集群加权降低了41.2个百分点[95%置信区间:27.7,54.9];任务完成率从79.1%降至60.9%,而配对的安全有用完成率上升了21.8个百分点[9.5,35.2]。部分答案重建了从未进入上下文的值或使用过滤后的行数作为预言机:塑造一次执行并非无干扰。写入控制、持久批准以及时间和聚合策略不在本次评估范围内。

英文摘要

Give an agent a human's credential and it inherits the person's reach without the judgment that limits its use. It can sweep every reachable record into model context, where hidden instructions steer its next call, and every request stays credential-valid while the agent exceeds its job or absorbs a secret. Prompts are a brittle guardrail: one fallible reasoner interprets the task and enforces its limits. We present Out-of-Band Policy Enforcement (OBPE), a trusted boundary outside agent reasoning. It authorizes the typed operation and resource, narrows the query before the backend call, then filters records and fields or masks values in the response. Semantic gating can deny or hold an authorized call on argument values or external state. A data policy owner sets the maximum grant; agent policy can only narrow it. We prove, under stated conditions, that the policy plan is order-independent and agent policy cannot widen the ceiling. Field removal covers one execution; masking and history rules claim less. We release an HTTP proxy prototype simplified from our production system, with conformance tests tying its typed Cedar policy core to the model. Against Jira and ServiceNow mocks, our benchmark compares prompted agents with and without OBPE on four models, including 20 adaptive red-team tasks. A trace failure means protected data entered agent context, an exact value appeared in the answer, or a forbidden effect completed. In 3,621 trials it fell from 57.6% to 0.2%, a cluster-weighted reduction of 41.2 points [95% CI: 27.7, 54.9]; fulfillment fell from 79.1% to 60.9%, while paired safe-useful completion rose 21.8 points [9.5, 35.2]. Some answers reconstructed a value that never entered context or used filtered row counts as an oracle: shaping one execution is not noninterference. Write controls, durable approval, and temporal and aggregate policies lie outside this evaluation.

发表机构

  • Redpanda Data(红熊猫数据公司)

机构由 AI 辅助整理,请以论文原文为准。

↑