arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18360cs.SEcs.AIcs.CY

仅一个门控不够:为智能体AI构建有状态的行动前控制机制

One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI

Gaston Besanson

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对智能体AI的多行动前控制耦合问题,提出修正-重门控协议,证明部分修正算子不可交换,经实验验证了相关机制的有效性。

中文摘要 AI 辅助

智能体AI系统执行重要行动时,需同时受多个行动前控制机制的约束,这些机制包括权限门控、资源门控和证据门控,它们可在行动执行前允许、降级或修正该行动。本文的核心研究对象是修正诱导的控制耦合:某一控制机制施加的修正会改变另一控制机制所评估的行动、证据或上下文,导致该控制机制此前的判断失效。我们对这种耦合进行了形式化,并提出了修正-重门控协议,该协议在给定假设下,可在当前有界、幂等的环境中恢复每个行动的合理性。我们进一步证明,两种已实现的修正算子(证据替换和资源预算下路由)不可交换——有限模型检查器找到了具体的反例实例——这使得修正顺序成为控制平面语义的一部分,而非实现细节。信任自身最近一次允许写入的受控证据缓冲区是状态层面同一问题的另一个实例:当前的允许性并不意味着未来参考的可信度,且易受未被发现的缺陷类别的投毒;两种缓解措施可减少但无法消除这种暴露。支撑结果确定了门控结果的正权重线性聚合可弥补成员否决的精确条件、统一的跨控制证据集,以及组合不会产生新的检测覆盖范围(已如实报告)。实验方面,在组合三个未修改的已发布引擎的确定性开放数据工件上,CH1-CH5在所有30个预注册种子下均符合其注册决策规则;CH6在W1工作流下符合,但在较小的W2工作流下不符合(已如实报告)。这是在带有合成元数据层的开放有效载荷数据上的机制演示,并非关于生产中普遍性的主张。

英文摘要

Agentic AI systems take consequential actions governed by more than one pre-action control at once: authority, resource, and evidence gates that can admit, degrade, or remediate an action before it executes. This paper's central object is remediation-induced control coupling: a remediation applied by one control can change the action, evidence, or context another control evaluates, invalidating that control's earlier judgment. We formalize this coupling and give a remediate-and-regate protocol that restores per-action soundness in the current bounded, idempotent setting under its stated assumptions. We further show that the two implemented remediation operators (evidence substitution and resource-budget downroute) do not commute -- a finite-model checker finds concrete counterexample instances -- making remediation order part of the control-plane semantics rather than an implementation detail. A governed evidence buffer that trusts its own most recent admitted write is a further instance of the same problem at the level of state -- current admissibility does not imply future reference trustworthiness -- and is vulnerable to poisoning from declared-uncovered defect classes; two mitigations reduce, not eliminate, that exposure. Supporting results establish the exact condition under which positive-weight linear aggregation of gate outcomes can compensate a member veto, a unified cross-control Evidence Set, and that composition manufactures no new detection coverage, reported honestly. Empirically, on a deterministic open-data artifact composing three published engines unmodified, CH1-CH5 meet their registered decision rules across all 30 pre-registered seeds; CH6 does so under W1 but not under the smaller W2 workflow, reported as such. This is a mechanism demonstration on open payload data with a synthetic metadata layer, not a claim about production prevalence.

发表机构

  • Universidad Torcuato Di Tella(托尔克瓦托·迪·特拉大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑