arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RLCascadeRouter:基于强化学习的无质量估计器级联路由

RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning

Shihong Huang, Shengjie Wang, Hong Ma, Zhou Xu

arXiv 2608.15817首次发表:更新:

发表机构

Polytechnic Institute, Zhejiang University; The Hong Kong Polytechnic University(浙江大学理工学院; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LLM路由的灵活性问题,提出RLCascadeRouter框架,将级联路由建模为马尔可夫决策过程,无需质量估计器,在LLMRouterBench基准上实现了更优的性能-成本权衡,且可纳入未见模型无需重训。

AI 中文摘要

大型语言模型(LLMs)不断发展的生态系统为优化性能-成本权衡提供了巨大潜力。然而,它们的异构能力和推理成本使得高效路由查询成为一项重大挑战。现有范式不够灵活:一次性路由器在观察响应前就做出承诺,而传统级联虽能自适应停止但遵循固定模型顺序。级联路由通过在每次响应后重新考虑是否停止或调用另一模型,消除了上述两种限制。当前方法采用“预测后优化”流程,估计响应质量和未来模型效用。但质量或效用的预测损失并不等同于路由决策损失:较低的预测误差不一定产生更好的动作,微小的边界交叉误差可能会反转“停止”或模型选择决策。因此,我们提出RLCascadeRouter,这是一个无质量估计器的框架,将级联路由表述为马尔可夫决策过程,动作包括“停止”和模型选择。它利用轨迹回报和优势直接优化性能-成本目标,其级联策略网络对模型选择的候选互补性和停止的剩余动作价值进行建模,消除了独立的事后响应质量估计器。在包含13个LLMs的10个LLMRouterBench基准测试中评估,RLCascadeRouter优于强大的基线,实现了更优的性能-成本权衡,无需重新训练即可纳入未见模型, ablation研究验证了两个策略组件的有效性。

英文摘要

The growing ecosystem of large language models (LLMs) offers huge potential to optimize performance-cost trade-offs. However, their heterogeneous capabilities and inference costs make efficiently routing queries a significant challenge. Existing paradigms are inflexible: one-shot routers commit before observing responses, whereas conventional cascades stop adaptively but follow a fixed model order. Cascade routing removes both restrictions by reconsidering whether to stop or invoke another model after each response. Current methods use a predict-then-optimize pipeline estimating response quality and future model utility. However, prediction loss for quality or utility is not equivalent to routing-decision loss. A lower prediction error does not necessarily yield a better action; a small boundary-crossing error can reverse a ``stop'' or model-selection decision. Therefore, we propose RLCascadeRouter, a quality-estimator-free framework that formulates cascade routing as a Markov decision process with actions comprising ``stop'' and model selection. It uses trajectory returns and advantages to directly optimize the performance-cost objective. Its Cascade Policy Network models candidate complementarity for model selection and remaining-action value for stopping, eliminating independent post-hoc response-quality estimators. Evaluated across ten LLMRouterBench benchmarks with thirteen LLMs, RLCascadeRouter outperforms strong baselines and achieves superior performance-cost trade-offs. It incorporates unseen models without retraining, and ablation studies validate both policy components.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑