人工智能代理的解释约束工具执行:无需信任模型原理的服务器验证动作声明
Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
- Accentrust(Accentrust公司)
- Georgia Institute of Technology(佐治亚理工学院)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对人工智能代理工具执行问题,提出EBTE中介层,将原理内容转换为动作声明并经服务器检查。通过形式化和实现参考配置文件,经多场景验证其可行性与诊断价值,支持服务器检查动作声明而非其他相关特性。
AI中文摘要:
使用工具的代理会暴露结构化调用,但通常会附加自由形式的原理。这些原理既不是授权也不是可靠的自省。我们提出了解释约束工具执行(EBTE),这是一个携带声明的中介层,它将与决策相关的原理内容转换为类型化的动作声明,并根据服务器持有的意图、策略、有效载荷、工具、风险、出处和新鲜度事实进行检查。EBTE不会扩大基线权限:冲突会拒绝,不完整或不确定的声明会被审查,只有匹配的声明才有资格进行受管执行。我们在显式中介和可信事实假设下对这种组合进行了形式化,并实现了一个具有最小化审核数据包的版本化参考配置文件。在136个编写的一致性场景中,完整配置文件匹配所有指定的处置,不接受96个指定的硬矛盾,通过232个变质检查;这些结果验证了所包含的配置文件,而不是总体性能。一个仅草案的参考集成在EBTE下不会转发48个编写的硬案例中的任何一个,同时保留所有16个软审查和4个对齐的草案路径。在一个冻结的2026年7月12日探索性224次尝试的托管模型记录中,历史生成/运行器协议计数分别为71/96、66/96和19/32;在当前管道下对保留的最小化声明进行单独标记的零调用事后重新验证产生了70/96、65/96和17/32。在一个源自AgentDojo的语义检查中,现有的高风险控制已经使所有12个攻击提议不被允许;EBTE还将它们解析为拒绝。这些结果支持了服务器检查动作声明的可行性和诊断价值,而不是原理的忠实性、人工审查的好处、代表性攻击抗性或生产安全性。
英文摘要:
Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that converts decision-relevant rationale content into typed action claims and checks them against server-held intent, policy, payload, tool, risk, provenance, and freshness facts. EBTE cannot widen baseline authority: conflicts deny, incomplete or uncertain claims review, and only matching claims remain eligible for governed execution. We formalize this composition under explicit mediation and trusted-fact assumptions and implement a versioned reference profile with minimized audit packets. Across 136 authored conformance scenarios, the full profile matches all specified dispositions, admits none of 96 designated hard contradictions, and passes 232 metamorphic checks. A draft-only reference integration forwards none of 48 authored hard cases under EBTE while preserving all 16 soft-review and 4 aligned draft paths. In a frozen 2026-07-12 exploratory 224-attempt hosted-model record, the historical generation/runner agreement counts are 71/96, 66/96, and 19/32; a zero-call revalidation of the preserved minimized claims under the current pipeline yields 70/96, 65/96, and 17/32. In an AgentDojo-derived semantic check, existing high-risk controls make all 12 attack proposals non-allow, while EBTE resolves the task--proposal contradictions as deny. Together, these studies establish profile conformance and demonstrate the feasibility of server-checked action claims within the evaluated settings.