批准洗白:系统化AI编码代理框架中的批准-执行绑定失败
Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对AI编码代理框架,提出“批准洗白”六类失败模式分类法,并原型化批准令牌以消除委托与时间洗白,揭示字段级验证的局限。
AI中文摘要:
现代AI编码代理框架(Claude Code、Codex CLI、Cursor)将其安全边界建立在一个很大程度上未经检验的假设上:即人类批准的操作为A与框架执行的操作为A'是相同的,其中A由既定策略确定,该策略规定了范围授权或会话级批准所授权的内容。我们证明这一假设会系统性地且可重复地失效。我们引入“批准洗白”这一概念,这是一种包含六种失败模式的分类法,通过这些模式,框架的强制机制在批准后静默地用A'替换A:范围洗白、参数洗白、时间洗白、工具洗白、委托洗白和语义洗白。与先前针对静态语料库评估风险分类器或推断隐式授权边界的工作不同,我们研究凭证绑定完整性:给定一个已批准的操作为,框架是否恰好分派该操作?通过对Claude Code的预执行中介点(PreToolUse)进行插桩,我们对所有六类进行了受控、无头、重复测量研究(每类N=19-20次运行),报告了带有Wilson置信区间的绑定缺口率(BGR)和评分者间一致性(kappa=1.0)。我们原型化了批准令牌,这是一种带密钥的能力Hk(主体, 代理ID, 会话ID, 工具, 参数, 范围, 过期时间),由中介方签发,且绝不将密钥返回给代理,通过118次运行的配对前后重放(McNemar精确检验)进行评估。该令牌完全消除了委托洗白,并且对于我们植入的会话身份不匹配构造,消除了时间洗白(p<10^-5),但按设计对范围洗白没有影响,并且对参数洗白没有显著减少(p=1):这是一个诚实的负面结果,因为这两类使每个记录的派发字段保持不变,在仅字段验证器可观察的进程级别以下发生分歧。我们讨论了仅在工具调用边界绑定的防御措施的影响。
英文摘要:
Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A' the harness executes, where A is fixed by a stated policy for what a scope grant or session-scoped approval authorizes. We show this assumption fails systematically and reproducibly. We introduce Approval Laundering, a taxonomy of six failure modes by which a harness's enforcement mechanism silently substitutes A' for A after approval: Scope, Argument, Temporal, Tool, Delegation, and Semantic laundering. Unlike prior work that evaluates risk classifiers against static corpora or infers implicit authorization boundaries, we study credential-binding integrity: given an already-approved action, does the harness dispatch exactly that action? Instrumenting Claude Code's pre-execution mediation point (PreToolUse), we conduct a controlled, headless, repeated-measures study of all six classes (N=19-20 runs each), reporting a Bound-Gap Rate (BGR) with Wilson confidence intervals and inter-rater agreement (kappa=1.0). We prototype Approval Token, a keyed capability Hk(principal, agent_id, session_id, tool, arguments, scope, expiry) issued by a mediator that never returns the key to the agent, evaluated via paired before/after replay of 118 runs (McNemar's exact test). The token fully eliminates Delegation laundering and, for our seeded session-identity-mismatch construction, Temporal laundering (p<10^-5), but by design leaves Scope laundering unaffected and shows no significant reduction in Argument laundering (p=1): an honest negative result, since these two classes leave every recorded dispatch field unchanged, diverging one process level below what a field-only verifier can observe. We discuss implications for defenses that bind only at the tool-invocation boundary.