谁来验证图?语言智能体因果动作验证中的错误设定攻击
Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents
浏览论文内容
中文总结 AI 辅助
针对因果动作验证器的错误设定攻击可导致高错误执行率,审计仅能防止错误行为,恢复安全与价值需额外实验成本。
中文摘要 AI 辅助
因果动作验证器通过检查每个提议的干预是否能够针对已提交的动作-状态图进行识别,来门控智能体的状态改变工具调用,并发出携带识别论证和单侧下置信界的证书。其中一个这样的验证器CIVeX,在一个混杂的工具使用基准上报告了零次错误执行。我们通过仅破坏已提交的图来对其进行红队测试。省略一条双向边,在基准公布的混杂强度下,使其从零错误执行变为15.3%的错误执行,其中91%的执行是有害的,效用从+2.27降至+0.35。反转一个箭头方向,使得一个中介被错误设定为混杂因子,导致48.9%的错误执行且没有正确执行。这些动作中的每一个都携带内部有效的证书。一个证明步骤,针对有界随机样本测试每个经观察认证的执行,检测到了这两种攻击,在真实图上555次执行中有2次误报;拒绝未通过测试或无法测试的执行,在我们测量的所有设置中均实现了零错误执行。它并未恢复有益执行:在公布的强度下,97.1%的有益动作仍然从未被执行,因为相同的错误设定在证明运行之前就拒绝了它们。那些拒绝也携带证书,审计它们有效,但其成本随拒绝次数而非执行次数扩展。恢复安全性每1,050个动作需花费127次实验;恢复损失的价值需再花费614次,此时经过审计的验证器在每个实例上做出诚实图的决策,并恰好用完其实验预算。仅检查执行的审计可防止错误行为。错误的不作为必须单独付费。
英文摘要
Causal action verifiers gate an agent's state-changing tool calls by checking whether each proposed intervention is identifiable against a committed action-state graph, and they issue a certificate that carries the identification argument and a one-sided lower confidence bound. One such verifier, CIVeX, reports zero false executions on a confounded tool-use benchmark. We red-team it by corrupting only the committed graph. Omitting a single bidirected edge takes it from zero false executions to 15.3% at the benchmark's published confounding strength, with 91% of its executions harmful and utility falling from +2.27 to +0.35. Reversing one arrowhead, so that a mediator is committed as a confounder, gives 48.9% false executions and no correct ones. Every one of these actions carries an internally valid certificate. An attestation step that tests each observationally certified execution against a bounded randomised sample detected both attacks, with 2 false alarms in 555 executions on a truthful graph; refusing what fails the test, or cannot be tested, gave zero false executions in every setting we measured. It does not restore beneficial execution: at the published strength 97.1% of beneficial actions are still never executed, because the same misspecification rejects them before attestation runs. Those rejections carry certificates too, and auditing them works, but its cost scales with the number of rejections rather than the number of executions. Recovering safety costs 127 experiments per 1,050 actions; recovering the lost value costs 614 more, at which point the audited verifier makes the honest graph's decisions on every instance and spends exactly its experiment budget. An audit that inspects only executions protects against wrongful action. Wrongful inaction has to be paid for separately.
发表机构
- The Tesseract Academy(泰瑟拉克学院)
机构由 AI 辅助整理,请以论文原文为准。