为智能体AI设计可靠的提交门控:共性数据故障下的成本感知验证组合
Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures
- Washington University in St. Louis(华盛顿大学圣路易斯分校)
- Southern Methodist University(南方卫理公会大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对智能体提交门控,提出VP-CONTROL框架,通过2×2实验和组合控制器,证明证据来源多样性比模型多样性更关键,并验证了原子防护与幂等标识符的有效性。
AI中文摘要:
智能体系统会执行改变状态的行动,但额外的验证器可能继承相同的上游故障。我们提出了VP-CONTROL,一种用于成本感知提交门控的运行时保障设计和确定性基准。其48个任务模板在六种故障机制下产生了2,880个场景。一个固定调用的2×2实验将验证器模型多样性与证据来源多样性区分开来。在来自两个本地行动者家族的冻结提案上,基于共享证据的跨模型投票批准了62.9%的不安全提案,而使用独立来源时仅为22.9%。来源效应为40.9个百分点,而模型多样性效应为11.3个百分点。一个组合控制器仅使用部署可观测的元数据来选择验证计划。在名义5%的每任务目标下,近似聚类调整校准在锁定测试上产生了1.9%的不安全执行率和38.2%的自动化安全覆盖率。匹配预算的组合也优于固定验证策略。迁移仍然是有条件的:未见过的故障家族产生16-26%的风险,而FinQA检查未能用测试的小型验证器重现来源效应。一项预注册的实时HTTP/SQLite研究测试了并发写入和丢失响应。检查后竞争击败了仅验证器的门控;事务性部分防护仅防止已覆盖的故障,而完整原子防护在216个回合中记录了零不安全效应。幂等请求标识符在丢失响应后防止了重复效应。结果激励了显式证据谱系、成本感知选择和提交时执行,同时揭示了近似校准和本地工具泛化的局限性。
英文摘要:
Agentic systems commit state-changing actions, but additional verifiers can inherit the same upstream fault. We present VP-CONTROL, a runtime-assurance design and deterministic benchmark for cost-aware commit gates. Its 48 task templates yield 2,880 scenarios across six fault regimes. A fixed-call 2 x 2 experiment separates verifier-model diversity from evidence-source diversity. On frozen proposals from two local actor families, a cross-model vote over shared evidence approves 62.9% of unsafe proposals, versus 22.9% with an independent source. The source effect is 40.9 percentage points, compared with 11.3 for model diversity. A portfolio controller selects verification plans using only deployment-observable metadata. Approximate cluster-adjusted calibration at a nominal 5% per-task target yields 1.9% unsafe execution and 38.2% automated safe coverage on the locked test. Matched-budget portfolios also improve on fixed verification policies. Transfer remains conditional: unseen fault families yield 16-26% risk, and a FinQA check fails to reproduce the source effect with the tested small verifiers. A preregistered live HTTP/SQLite study tests concurrent writes and lost responses. After-check races defeat verifier-only gates; transactional partial guards prevent only covered failures, while a full atomic guard records no unsafe effects across 216 episodes. Idempotent request identifiers prevent duplicate effects after lost responses. The results motivate explicit evidence lineage, cost-aware selection, and commit-time enforcement, while exposing the limits of approximate calibration and local-tool generalization.