CAGE:面向使用工具的智能体的带类型返回不确定性的可验证授权
CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents
浏览论文内容
中文总结 AI 辅助
该研究针对使用工具的LLM智能体的授权问题,提出CAGE方法直接验证带绑定错误和数值漂移的联合邻域,可消除点式网关的预算内错误授权,保留部分自主决策,适配多种场景。
中文摘要 AI 辅助
使用工具的大语言模型(LLM)智能体基于带类型的工具返回采取行动,这类返回记录了来源信息以及带有数值的分类字段。运行时权限网关通常会对观察到的返回和操作进行授权,这使得决策无法抵御返回与其来源绑定过程中出现的小错误。我们研究的问题是:在声明的合理绑定返回邻域内,候选操作是否仍保持授权,该邻域包含一个可接受的绑定错误以及有界的数值漂移。我们证明,单独验证分类通道和数值通道无法组合:仅在单个通道上安全的扰动,共同作用可能会使同一操作变得不安全。CAGE 直接验证该联合邻域,精确枚举离散分支并在每个分支内验证连续扰动。在合成、基于代码的策略、监管以及真实交易场景中,CAGE 消除了精确点式网关所允许的预算内错误授权,同时保留了相当一部分自主决策。当策略可执行时,CAGE-Exact 直接验证策略本身;否则,CAGE-Lip 和 CAGE-RS 在明确的、经测量的保真度假设下验证学习到的网关。
英文摘要
Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed return and action, leaving the decision unprotected against small errors in how the return was bound to its source. We ask whether a candidate action stays authorized over a declared neighborhood of plausible correctly bound returns: one admissible binding fault plus bounded numerical drift. We prove that certifying the categorical and numerical channels separately does not compose: perturbations that are safe on each channel alone can jointly turn the same action unsafe. CAGE certifies this joint neighborhood directly, enumerating the discrete branches exactly and certifying the continuous perturbation within each branch. Across synthetic, policy-as-code, regulatory, and real-transaction settings, CAGE removes the in-budget false allows that accurate pointwise gates admit, while keeping a useful fraction of decisions autonomous. When the policy is executable, CAGE-Exact certifies the policy itself; otherwise CAGE-Lip and CAGE-RS certify a learned gate under an explicit, measured fidelity assumption.