从版本冲突到决策冲突:面向长时间运行AI智能体的选择性重新验证
From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents
- Huawei Technologies Co., Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对长时间运行AI智能体因状态变化导致操作失效的问题,提出ATR方法,通过记录并选择性重验决策条件,在受控实验中实现零错误且高效验证。
AI中文摘要:
长时间运行的AI智能体可能会读取状态、进行推理、等待工具或人工审批,并在很久之后才执行外部操作。在此期间,原本支撑该操作合理性的状态可能已经发生变化。例如,在智能体根据100英镑的限额提出80英镑退款后,客户姓名的变更仅影响展示元数据,新的90英镑限额仍然允许该退款,而50英镑的限额则使其无效,另一名工作人员已发放的退款则必须防止重复操作。标准的乐观并发控制和版本检查能够检测到先前读取的状态已发生变化,但本身并不能确定这种变化是否会使待执行操作的合理性失效。我们将任何检测到的版本变化称为版本冲突;当该变化使操作的合理性失效时,它同时也是决策冲突。ATR记录显式、可执行的待执行操作合理性条件,并在释放外部操作前仅重新检查受变化影响的条件。它可以保留操作、刷新非决定性元数据、要求重新规划或阻止执行;目标端事务或比较并交换操作将已检查的状态绑定到提交。在15个变异案例的210,000次受控执行中,ATR匹配了所有开发者指定的结果,没有出现错误的允许或阻止。在十个持久的SQLite检查点/恢复单元中,它每次变化评估0.6个条件,而FullScan为6.0个。在4,093次记录读取时,ATR耗时9.3微秒,而FullScan为2595.9微秒。这些确定性结果确立了受控可行性,而非生产通用性或所需条件的自动提取。
英文摘要:
Long-running AI agents may read state, reason, wait for tools or human approval, and perform an external action much later. The state that justified the action can change in the meantime. For example, after an agent proposes an 80 GBP refund under a limit of 100, a customer-name change affects only presentation metadata, a new limit of 90 still permits the refund, a limit of 50 invalidates it, and a refund issued by another worker must prevent a duplicate. Standard optimistic concurrency control and version checks can detect that previously read state has changed, but by themselves do not determine whether that change invalidates the pending action's justification. We call any detected version change a version conflict; when that change invalidates the action's justification, it is also a decision conflict. ATR records the explicit, executable conditions that justify a pending action and rechecks only the conditions affected by a change before releasing the external operation. It can retain the action, refresh non-decisive metadata, require replanning, or block execution; a target-side transaction or compare-and-set binds checked state to commit. Across 210,000 controlled executions over 15 mutation cases, ATR matched every developer-specified outcome with no false allows or blocks. In ten durable SQLite checkpoint/resume cells, it evaluated 0.6 conditions per change versus 6.0 for FullScan. At 4,093 recorded reads, ATR took 9.3 microseconds versus 2595.9 microseconds for FullScan. These deterministic results establish controlled feasibility, not production generality or automatic extraction of the required conditions.