arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22951cs.AIcs.CLcs.LGcs.MA

AgentRouter:面向成本最优多步智能体工作流的异构模型路由

AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows

Rudrendu Kumar Paul, Sourav Nandy

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体工作流中步骤复杂度差异大的问题,提出轻量级分类器AgentRouter,将轨迹步骤路由至四层模型,实现72%成本降低并保留97.3%质量。

中文摘要 AI 辅助

企业级智能体系统若将每个轨迹步骤都路由至前沿模型,会在子任务上浪费60-80%的推理预算,而这些子任务由较小模型处理同样出色。现有路由方案优化单轮查询分配,却忽略了智能体工作流所独有的特性:单个轨迹内的子任务复杂度差异巨大。规划步骤可能需要前沿级推理,而随后的格式化步骤仅需7B模型即可。我们将步骤级模型路由形式化为智能体轨迹上的序列分配问题,并提出AgentRouter——一个轻量级分类器(12M参数,在A100 GPU上每步开销小于5ms),利用路由时可提取的五个特征,将每个轨迹步骤映射到四个模型层级之一。AgentRouter在涵盖规划、编码、研究和数据分析任务的50,000个标注智能体轨迹步骤上训练,相对于仅使用前沿模型的基线实现了72%的成本降低,同时保留了前沿模型97.3%的质量(端到端任务完成率下降不到3%);在最小复杂度步骤上,逐步路由准确率达到91%,在高效层级步骤上达到85%,而在较难的中间层级和前沿层级上为76-82%。在相同基准上,RouteLLM和FrugalGPT(按步骤应用)仅分别实现31%和44%的成本降低,因为它们的单轮训练信号忽略了轨迹级质量依赖关系。

英文摘要

Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller models handle equally well. Existing routing solutions optimize single-turn query assignment but ignore a property unique to agentic workflows: subtask complexity varies widely within a single trajectory. A planning step may require frontier-class reasoning while a subsequent formatting step needs only a 7B model. We formalize step-level model routing as a sequential assignment problem over agent trajectories and propose AgentRouter, a lightweight classifier (12M parameters, <5ms overhead per step on an A100 GPU) that maps each trajectory step to one of four model tiers using five features extractable at routing time. Trained on 50,000 annotated agent trajectory steps spanning planning, coding, research, and data analysis tasks, AgentRouter achieves 72% cost reduction relative to frontier-only baselines, retaining 97.3% of frontier-only quality (less than 3% degradation in end-to-end task completion); per-step routing accuracy reaches 91% on minimal-complexity steps and 85% on efficient-tier steps, with 76-82% on the harder mid-range and frontier tiers. On the same benchmarks, RouteLLM and FrugalGPT (applied per-step) achieve only 31% and 44% cost reduction respectively, because their single-turn training signal misses trajectory-level quality dependencies.

发表机构

  • Boston University(波士顿大学)
  • University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑