RSI-Router:面向成本高效智能体的子任务级LLM路由与技能进化
RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents
浏览论文内容
中文总结 AI 辅助
RSI-router通过递归自我改进构建子任务级模型路由与技能,在五个智能体基准上以约一半推理成本超越强基线,并建立更优性能-成本帕累托前沿。
中文摘要 AI 辅助
大型语言模型(LLM)智能体的实际部署需要在可承受的推理成本下实现强大的任务性能。对于长时程智能体任务,可以通过任务内大-小模型协作来改善这种性能-成本权衡,因为较小的模型即使无法解决完整任务,也能处理某些阶段。本文提出了RSI-router,一种通过递归自我改进在累积经验上构建子任务级模型分配和模型特定技能的路由框架。每次迭代包含四个阶段:子任务挖掘从训练轨迹中推导子任务定义和识别规则;路由策略进化提出并评估多样化的模型分配;模型特定技能进化比较路由轨迹与仅大模型轨迹以诊断失败并开发可复用的执行技能;帕累托最优路由器选择使用历史和新生成的路由器更新帕累托种群,同时保留被支配的路由器作为后续进化的经验。在DeepSeek-V4.1-Flash和Qwen3.5-9B之间进行路由时,RSI-router在五个智能体基准上以约一半的推理成本(48.3%)持续超越仅使用DeepSeek的基线。特别是在ALFWorld、ScienceWorld和WebShop上,它将推理成本降低了74.7-82.2%,同时提升了性能;在Terminal-Bench 2.0上,它以18.0%的成本降低实现了16.7%的相对性能提升。此外,RSI-router建立了比9种路由方法更强的性能-成本帕累托前沿。
英文摘要
Practical deployment of large language model (LLM) agents requires strong task performance at affordable inference cost. For long-horizon agentic tasks, this performance-cost trade-off can be improved through within-task large-small model collaboration, as smaller models can handle some stages even when they cannot solve the full task. In this paper, we introduce RSI-router, a routing framework that constructs subtask-level model assignments and model-specific skills through recursive self-improvement over accumulated experience. Each iteration consists of four stages: Subtask Mining derives subtask definitions and identification rules from training trajectories; Routing Strategy Evolution proposes and evaluates diverse model assignments; Model-Specific Skill Evolution compares routed and large-model-only trajectories to diagnose failures and develop reusable execution skills; and Pareto-Optimal Router Selection updates the Pareto population using historical and newly generated routers while retaining dominated routers as experience for subsequent evolution. Routing between DeepSeek-V4.1-Flash and Qwen3.5-9B, RSI-router consistently surpasses the DeepSeek-only baseline at roughly half the inference cost (48.3%) across five agentic benchmarks. In particular, on ALFWorld, ScienceWorld, and WebShop, it cuts inference cost by 74.7-82.2% while simultaneously improving performance; on Terminal-Bench 2.0, it achieves a 16.7% relative performance gain at 18.0% lower cost. Moreover, RSI-router establishes a stronger performance--cost Pareto frontier than 9 routing methods.
发表机构
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。