发表机构
The Hong Kong University of Science and Technology (Guangzhou); Singapore University of Technology and Design(香港科技大学(广州); 新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出多智能体系统执行一致性定义,通过证据条件区分相同记录下的合规与违规执行,并借助CAVERT框架在诊断与恢复中优于基线,强调接口需保留执行证据。
AI 中文摘要
共享内存协调智能体的行动,但正确的记录并不能证明这些行动满足任务要求。内存治理和故障诊断规范或检查记录信息;它们本身并不能确定这些信息是否足以判断任务职责。我们通过管理状态使用、信息交接和最终状态一致的职责来定义执行一致性,并给出判断履行的明确证据条件。我们的核心主张是,在相同的任务规则下,相同的保留记录可能对应合规执行和违规执行。受控移除诸如回执、行动依赖或响应有效性等证据,会使82.4%的相反标签对无法区分;恢复证据则能分离97.9%的合并对。自然日志注释识别了实际执行中定义的违规行为。然而,现有日志并不总是明确表示这些判断所需的执行关系。为评估该定义的实际价值,我们使用CAVERT(一致性诊断与恢复框架)从日志中提取支持的关系并应用这些标准。在所有12个基准执行器设置中,它在诊断方面始终优于合同提示的LLM和基于规则的基线。在相同的门控和执行器限制下,它在所有四个评估环境中的规则引导恢复方面也表现更优。这些发现确定了智能体内存和执行接口应保留的执行证据,以便进行可靠判断。
英文摘要
Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded information; they do not by themselves establish whether it is sufficient to judge task duties. We define execution consistency through duties governing state use, information handoffs, and final-state agreement, with explicit evidence conditions for judging fulfillment. Our core claim is that identical retained records can correspond to compliant and violating executions under the same task rule. Controlled removal of evidence such as receipt, action dependence, or response validity leaves 82.4% of opposite-label pairs indistinguishable; restoration separates 97.9% of the merged pairs. Natural-log annotations identify the defined violations in actual executions. However, existing logs do not always explicitly represent the execution relationships needed for these judgments. To assess the definition's practical value, we use CAVERT, a framework for consistency diagnosis and recovery, to extract supported relationships from logs and apply these criteria. It consistently outperforms contract-prompted LLM and rule-based baselines in diagnosis across all 12 benchmark-executor settings. Under the same gate and executor limits, it also outperforms rule-guided recovery in all four evaluated environments. These findings identify execution evidence that agent-memory and execution interfaces should preserve for reliable judgment.
Comments39 pages, 7 figures, 30 tables (including appendix)