智能体网格:非幂等智能体委托的可靠性原语——身份充分性与证据充分性
Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy
浏览论文内容
中文总结 AI 辅助
本文研究生产智能体平台的故障问题,发现现有服务网格机制的假设不成立,提出身份与证据充分性,推导7个以委托为单元的可靠性原语。
中文摘要 AI 辅助
自主智能体在编排器的协调下日益执行有界软件任务,编排器采用服务网格的机制(如重试、超时和错误率熔断)来处理这些任务。我们对一个生产级智能体软件交付平台开展故障研究,涉及147起编号事件、81次运行,每次运行都有可测量的成本,且多数情况配有可复现故障的变异证明。这些机制所依赖的全部三个假设在实践中均被违反,我们量化了相关后果:54次连续成功的工具调用,任何错误率熔断都无法察觉;进度信号因自身构造而恒定,保证在第三次修复轮次触发误跳闸,导致6个组件的一次运行缩减至3个;一次委托的6次调用累计产生21个事件,使正确的幂等组件无法推进;故障路由错误导致5个组件为2个组件的故障被唤醒,留下3个旁观者回退正常工作代码;12起事件中,执行层阻止了正确工作,其中最昂贵的一次消耗107次智能体轮次且无任何写入被接受。我们发现一个跨领域原因及其对偶:身份充分性——5个独立子系统中,无法区分的身份产生了自信的错误答案,其中2个子系统独立推导了修正规则;证据充分性——可靠性决策只能基于能够移动、可归因于所测量对象且在相同条件下具有确定性的证据。基于这些发现,我们推导了7个可靠性原语,其执行单元是委托而非消息,并指定了本研究所推动但未构成的受控评估。
英文摘要
Autonomous agents increasingly perform bounded software tasks under an orchestrator that retries, resumes, and budgets them. The machinery such orchestrators reach for is the service mesh's: retry, timeout, and error-rate circuit breaking. We report a failure study of a production agentic software-delivery platform over 147 numbered incidents spanning 81 runs, each with a measured cost and, in most cases, a mutation proof reproducing the failure. All three assumptions those primitives rest on are violated in practice, and we quantify the consequences: a loop of fifty-four consecutive successful tool calls no error-rate breaker could see; a progress signal constant by construction, guaranteeing a false trip on the third repair round and driving one run from six of six components to three; twenty-one events accumulated across six invocations of one delegation, making a correct, idempotent component unwinnable; a misrouted failure that woke five components for a two-component fault, leaving three bystanders regressing working code; and twelve incidents in which the enforcement layer blocked correct work, the most expensive costing 107 agent turns and zero accepted writes. We find one cross-cutting cause and its dual. Identity adequacy: in five separate subsystems an identity that failed to discriminate produced a confident wrong answer, and two of them derived the corrective rule independently. Evidence adequacy: a reliability decision may be taken only on evidence capable of moving, attributable to what it measures, and deterministic under identical conditions. From the findings we derive seven reliability primitives whose enforcement unit is the delegation rather than the message, and specify the controlled evaluation the study motivates but does not constitute.