OrchBench:通过确定性模拟单独评估多智能体编排计划
OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
浏览论文内容
中文总结 AI 辅助
研究多智能体编排计划评估问题,提出OrchBench基准,通过构建有向无环图及确定性模拟单独评估编排计划,模拟分数与Claude Code执行高度相关且资源需求少,为比较和诊断编排计划提供有效基准。
中文摘要 AI 辅助
复杂任务常分解为可并行但相互依赖的子任务,编排对多智能体系统性能至关重要。现有评估依赖端到端执行,混淆编排计划质量与工作能力等。我们提出OrchBench,一种基于模拟的基准,用于单独评估多智能体编排计划。它从现实任务构建有向无环图,给定相关条件后,评估规划器分配子任务并指定跨智能体信息传输等。确定性模拟器评估计划并返回可解释的结果质量等指标。模拟分数与Claude Code执行的质量分数高度相关,同时所需资源少。结果表明OrchBench是比较和诊断多智能体编排计划的有效且可解释的基准。
英文摘要
Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS). Existing evaluations typically rely on end-to-end execution, which conflates orchestration-plan quality with worker capabilities, tool reliability, and environmental noise. Moreover, the time and token costs of real execution grow rapidly with workflow scale, making systematic evaluation expensive. We present OrchBench, a simulation-based benchmark for evaluating multi-agent orchestration plans in isolation. Starting from real-world tasks, OrchBench constructs directed acyclic graphs (DAGs) that encode task dependencies, with controlled sizes and degrees of parallelism. Given a DAG, a per-agent context limit, and an agent budget, the evaluated planner assigns subtasks to agents and specifies cross-agent information transfers and their retention ratios. A deterministic simulator evaluates the resulting plan without invoking worker agents and returns interpretable measures of result quality, makespan, and token cost. The simulated scores produced by OrchBench correlate strongly with quality scores from Claude Code executions, achieving a Pearson correlation of \(r=0.816\), while requiring only \(1.3\%\) of the tokens and \(10.3\%\) of the wall-clock time. Across diverse planners and workflow scales, we find that preserving task-critical information is more important than simply increasing the number of agents, and the benefits of parallelism diminish as coordination failures accumulate. These results establish OrchBench as an efficient and interpretable benchmark for comparing and diagnosing multi-agent orchestration plans.