arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OrchestraBench:评估多智能体编排的故障模式、恢复能力与分解质量

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality

Yidian Chen, Yingzi Gu, Natan Vidra, Spurthi Setty, Sharon Zheng

arXiv 2608.05263首次发表:更新:

AI 中文总结

OrchestraBench 是评估多智能体编排故障模式、恢复能力与分解质量的基准,通过可控故障注入揭示了不同路由策略、智能体模式的故障处理差异及级联规律。

AI 中文摘要

多智能体编排框架正从演示阶段走向生产应用,但现有基准通常仅报告任务准确率,未诊断流水线为何失败、级联故障从何处开始,或是哪个路由决策导致了故障。OrchestraBench 是一个可控、可通过种子复现的故障注入工具,针对模板化企业工作流评估故障、恢复及分解质量,它引入级联半径和单故障模式恢复作为核心指标,并通过自举置信区间与配对检验对比路由策略。在含 26 个案例的金标准诊断任务中,关键字/标志路由器在具有误导性或缺失表面标志的对抗性案例中得分为 0%,而意图推理模型路由器得分为 100%,与 oracle(基准模型)表现一致。针对真实 Claude 智能体在可验证算术依赖链上的可控机制探测,揭示了五种 MAST 模式下的三级故障处理效果:工具故障完全恢复(1.0),委托歧义部分恢复(0.30),三种潜在或语义模式完全无法恢复(0.0)。当将计算重构为贷款审批工作流,并在 Sonnet、Opus、Haiku 模型上测试时,该排序依然存在,尽管绝对恢复率随上下文变化。盲目重试会重现潜在故障并增加检测时间,表明检测与归因对于故障遏制是必要的。级联半径随流水线深度增加而增大(深度 3-7 时均值为 0.9 至 4.7)。可信状态修复的 ablation(消融)实验显示,表观遏制增益主要来自可信状态信号,而非自主检测。这些结果是可控链机制探测的结论,并非针对领域工作负载的断言。

英文摘要

Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade began, or which routing decision caused the breakdown. OrchestraBench evaluates failure, recovery, and decomposition through a controlled, seed-reproducible failure-injection harness over templated enterprise workflows. It introduces cascade radius and per-failure-mode recovery as primary metrics and compares routing policies with bootstrap confidence intervals and paired tests. On a 26-case gold-labelled diagnostic, a keyword/flag router scored 0% on adversarial cases with misleading or missing surface flags, whereas an intent-reasoning model router scored 100%, matching the oracle. Controlled mechanism probes with a real Claude agent over a verifiable arithmetic dependency chain revealed three failure-handling tiers across five MAST modes: tool faults recovered fully (1.0), ambiguous delegation recovered partially (0.30), and three latent or semantic modes never recovered (0.0). This ordering persisted when the computation was reframed as a loan-approval workflow and across Sonnet, Opus, and Haiku, although absolute rates shifted with context. Blind retry reproduced latent faults and increased time to detection, indicating that detection and attribution are necessary for containment. Cascade radius increased with pipeline depth (mean 0.9 to 4.7 across depths 3-7). A trusted-state repair ablation showed that apparent containment gains primarily came from the trusted-state signal rather than autonomous detection. These results are controlled-chain mechanism probes, not domain-workload claims.

Comments8 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑