arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体审批洗白:超出已批准调用的传递效应

Agent Approval Laundering: Transitive Effects Beyond the Approved Invocation

Jinqian Zhang, Haojun Xia, Shujiang Wu, Jingkun Yue, Xia Zhang, Zhangpei Cheng, Bibo Tu

arXiv 2609.28586首次发表:更新:

发表机构

Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences; Beihang University; State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(中国科学院信息工程研究所; 中国科学院大学网络空间安全学院; 北京航空航天大学; 北京邮电大学网络与交换技术国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对编码智能体审批界面记录覆盖失败问题,提出审批洗白概念,形式化闭包绑定审批并构建基准,验证了记录仅策略的信息极限,通过效应绑定预测将残余效应从10降至3。

AI 中文摘要

编码智能体的审批界面将人类决策绑定到命令或工具调用,而开发者工具则执行该调用所激活的传递工作流。软件包安装可运行生命周期钩子并写入文件;MCP调用可行使网络权限。我们将由此产生的记录覆盖失败称为“审批洗白”:持久记录仅记载入口调用,却遗漏了其工作流所行使的效应。我们首次对智能体系统中这种记录到闭包的关系进行了系统性安全分析。我们形式化了六类效应上的闭包绑定审批,并推导出一个信息极限:相同的策略可见字段可能要求不同的效应特定决策,因此任何仅基于记录的策略都无法同时保证两者。审批到行动安全基准将审批对象和决策时元数据绑定到执行后证据。在111个固定的审批对象/轨迹对中,残余记录从显式字段下的40个降至带命令语义的17个,再降至带决策时元数据的13个。在11个固定SHA执行中,阶梯达到零元数据残余;三种产品前端中重复出现两种精确映射。对于前瞻性恢复,效应绑定记录在授权前提交冻结的、源支持的预测和出处。在17个预指定的保留工作流上,预测达到0.926的宏召回率和0.941的宏精确率;绑定它们将残余效应从10个削减至3个。Claude Code PreToolUse集成将冻结记录通过权限路径传递,无需自动审批。这些结果确立了审批洗白是一种可测量、可复现的记录覆盖失败,尽管调用身份真实。它们促使在授权前将每次调用绑定到其工作流传递效应边界的源支持预测,并将该绑定与决策一同保留。

英文摘要

Coding-agent approval interfaces bind a human decision to a command or tool call, while developer tools execute the transitive workflow that invocation activates. Package installation can run lifecycle hooks and write files; an MCP call can exercise network authority. We call the resulting record-coverage failure approval laundering: the durable record names the entry invocation but omits effects exercised by its workflow. We present the first systematic security analysis of this record-to-closure relation in agent systems. We formalize closure-bound approval over six effect classes and derive an information limit: identical policy-visible fields can require different effect-specific decisions, so no record-only policy can guarantee both. The Approval-to-Action Security Benchmark binds approval objects and decision-time metadata to post-execution evidence. Across 111 fixed approval-object/trace pairs, residual records fall from 40 under explicit fields to 17 with command semantics and 13 with decision-time metadata. Across 11 fixed-SHA executions, the ladder reaches zero metadata residuals; two exact mappings recur across three product frontends. For prospective recovery, effect-bound records commit frozen, source-backed predictions and provenance before authorization. On 17 prespecified holdout workflows, predictions achieve 0.926 macro recall and 0.941 macro precision; binding them cuts residual effects from 10 to 3. A Claude Code PreToolUse integration carries the frozen record through the permission path without automatic approval. These results establish approval laundering as a measurable, recurrent record-coverage failure despite truthful invocation identity. They motivate binding each invocation before authorization to a source-backed prediction of its workflow's transitive effect boundary and preserving that binding with the decision.

Comments19 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑