arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32917cs.AIcs.CLcs.MA

Planner-as-Router:面向成本高效多智能体工作流的规划时模型路由

Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows

Vivek Kumar Singh, Preeti Priyam, Gautam Bhowmick

首次发表
浏览论文内容

中文总结 AI 辅助

PaR将模型层级选择融入规划阶段,为多智能体工作流实现成本高效路由,在EntBench基准上保持成本-准确性前沿,成本较全前沿路由降低44%。

中文摘要 AI 辅助

在生产环境中运行大型语言模型(LLM)智能体会迅速变得昂贵。前沿模型(规模最大、能力最强的层级)虽然准确,但其每令牌成本可能是小模型的25倍,而且一旦工作流将多次调用串联起来,这种差距会进一步扩大。Planner-as-Router(PaR)从不同角度解决这一问题。它不将模型层级选择留给下游某个组件,而是将选择融入规划本身。当规划器将查询分解为子任务时,它同时为每个子任务分配一个模型规模层级(小、中或前沿,按能力和价格排序),从而在任何专家运行之前,子任务之间的依赖关系就已可见。与逐调用路由器(如级联路由)不同,后者一次只查看一个节点,PaR能预先看到整个工作流,且无需单独的路由器模型或训练数据。我们使用EntBench评估PaR,该基准包含七个类别的54个企业智能体任务,通过实际运行生成的SQL(结构化查询语言)和MongoDB查询对实时数据库进行评分。在跨越八个路由器和三个种子的1,157次评估中,PaR保持在观察到的成本-准确性前沿上。它在准确性上与一种汇合前沿启发式(仅对终端节点使用前沿模型)相当,成本相近;与忠实的FrugalGPT级联相比,成本更低;与全前沿路由相比,成本降低44%,同时仅牺牲2.9个百分点的准确性。部分准确性差距落在54任务研究的正负六个百分点置信区间内,因此我们将PaR的优势定位为前沿位置,而非明确的准确性胜利。我们还报告了一项初步观察,而非验证结果:一个小型试点提示,廉价路由可能在组合工作流上带来隐藏的复合惩罚,我们将其作为未来测量的假设。PaR、EntBench及所有评估代码均已开源。

英文摘要

Running large language model (LLM) agents in production gets expensive fast. A frontier model (the largest, most capable tier) is accurate but can cost 25 times what a small model costs per token, and the gap compounds once a workflow chains several calls together. Planner-as-Router (PaR) attacks this from a different angle. Instead of leaving model-tier selection to some component downstream, it folds the choice into planning itself. As the planner breaks a query into subtasks, it also assigns each one a model size tier (small, mid, or frontier, ordered by capability and price), so the dependencies between subtasks are visible before any specialist runs. Unlike per-call routers such as cascade routing, which look at one node at a time, PaR sees the whole workflow up front and needs no separate router model or training data. We evaluate PaR with EntBench, a benchmark of 54 enterprise agentic tasks across seven classes, graded by actually running the generated Structured Query Language (SQL) and MongoDB queries against live databases. Over 1,157 evaluations spanning eight routers and three seeds, PaR stays on the observed cost-accuracy frontier. It matches a sink-frontier heuristic (frontier model on terminal nodes only) in accuracy at comparable cost and a faithful FrugalGPT cascade at lower cost, and cuts cost 44% against all-frontier routing while giving up 2.9 points of accuracy. Several accuracy gaps fall inside the plus-or-minus six-point confidence interval of a 54-task study, so we frame PaR's advantage as frontier position rather than a clean accuracy win. We also report a preliminary observation, not a validated result: a small pilot hints that cheap routing may carry a hidden compounding penalty on compositional workflows, which we frame as a hypothesis for future measurement. PaR, EntBench, and all evaluation code are open source.

补充信息

↑