WISERouter:具有工作负载预算约束的大语言模型路由
WISERouter: LLM Routing with Workload Budget Constraint
- The University of British Columbia(英属哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对大规模使用LLMs成本高的问题,将LLM路由建模为约束上下文多臂老虎机问题,提出WISERouter框架,支持离线和在线学习。证明WR-Online有次线性遗憾界,实验表明WR-Offline性能超基线且更守预算,WR-Online用较少探索数据达可比性能。
AI中文摘要:
大语言模型(LLMs)在多个领域表现出色,但大规模使用最强大的模型成本过高。LLM路由通过为每个查询分配合适的模型来平衡效用和预算,利用模型能力和成本的多样性。当前方法存在局限性:要么使用不总是强制执行预算约束的启发式方法,要么采用固定的每个查询预算,无法适应工作负载且导致性能次优;还需要在密集数据集上进行监督学习,数据收集成本高。为应对这些挑战,我们将LLM路由表述为约束上下文多臂老虎机问题,并引入WISERouter框架,支持从历史交互中进行离线学习以及带有探索的在线学习。我们进一步证明WR-Online在时间范围T上实现了$O(\sqrt{T})$的次线性遗憾界。在RouterBench和SWE-Bench上的实证结果表明,WR-Offline在固定预算下性能超过现有基线且更紧密地遵守预算约束,WR-Online在使用少得多的探索数据的情况下实现了与基线相当的性能。
英文摘要:
Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale. LLM routing exploits diversity in model capability and cost by assigning each query to a suitable model to balance utility and budget. Current methods have two limitations: (i) they either use heuristics that do not always enforce the budget constraint or impose a fixed per-query budget that cannot adapt across the workload and leads to suboptimal performance; (ii) they require supervised learning on a dense dataset with statistics for every query-model pair, which is expensive to collect. To address these challenges, we formulate LLM routing as a constrained contextual multi-armed bandit problem and introduce WISERouter (WR for short), a framework that supports offline learning from historical interactions as well as online learning with exploration. We further prove that WR-Online achieves a sublinear regret bound of $O(\sqrt{T})$ over a time horizon $T$. Empirical results on RouterBench and SWE-Bench demonstrate that (i) WR-Offline surpasses existing baselines in performance under a fixed budget and adheres more closely to budget constraints, and (ii) WR-Online achieves comparable performance to the baselines, while using substantially less exploration data.