arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00107cs.LGcs.AI

MetaRoute-Bench:评估智能体工作流路由的元决策策略

MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflow Routing

Natan Vidra, Alina Kapanova, Arun Kanhai, Spurthi Setty

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出MetaRoute-Bench框架,构建含180个任务的基准,对比8种路由策略,发现任务感知组合策略成功率优于静态策略,明确路由组合与验证的关键作用,为智能体工作流路由评估提供可复现方法。

中文摘要 AI 辅助

智能体系统必须反复决定是直接回答、分解任务、调用工具、执行代码、委派给专家、验证中间结果,还是从失败中恢复。这些元决策不仅影响任务成功,还影响运营成本和延迟,但它们通常嵌入在编排框架中,仅通过总体任务准确率进行评估。我们提出MetaRoute-Bench,这是一个开放、可检查的框架,用于在共享执行模型下比较元决策策略。初始基准包含180个合成任务配置文件,涵盖数据分析、研究和文档处理,8种路由策略,以及30对随机种子。在43200条轨迹中,一种感知任务的组合策略达到79.4%的成功率,而强大的特定工作负载静态策略为76.7%,一次性任务路由为67.4%,直接回答为52.9%。相对于静态策略,这是2.7个百分点的提升,配对95%置信区间为±2.0个百分点,同时平均成本提高4.7%,延迟提高6.4%。 ablation实验显示,当路由组合限制为单一操作且移除验证时,损失最大。这些结果由带种子的离线执行模型生成,而非实时部署;因此,主要贡献是一种可复现的评估方法和对路由策略权衡的分析,而非生产有效性的证据。我们发布任务生成、策略、轨迹、测试和分析工件,以支持实时系统验证。

英文摘要

Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recover from failure. These meta-decisions affect not only task success but also operating cost and latency, yet they are often embedded inside an orchestration framework and evaluated only through aggregate task accuracy. We present MetaRoute-Bench, an open, inspectable framework for comparing meta-decision policies under a shared execution model. The initial benchmark contains 180 synthetic task profiles spanning data analysis, research, and document processing, eight routing policies, and 30 paired random seeds. Across 43,200 traces, a task-aware compositional policy achieves 79.4% success compared with 76.7% for a strong workload-specific static policy, 67.4% for one-shot task routing, and 52.9% for direct answering. Relative to the static policy, this is a 2.7 percentage-point improvement with paired 95% CI of plus or minus 2.0 points, at 4.7% higher mean cost and 6.4% higher latency. Ablations show the largest losses when route composition is restricted to one operation and when verification is removed. These results are generated by a seeded offline execution model rather than a live deployment; accordingly, the primary contribution is a reproducible evaluation method and an analysis of routing-policy tradeoffs, not evidence of production effectiveness. We release task generation, policies, traces, tests, and analysis artifacts to support live-system validation.

发表机构

  • Anote AI(阿诺特人工智能公司)
  • Cornell University(康奈尔大学)
  • CUNY(纽约城市大学)
  • Stevens Institute of Technology(史蒂文斯理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑