FlexRouter:学习互补模型集以实现灵活的LLM路由
FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing
浏览论文内容
中文总结 AI 辅助
FlexRouter通过行列式点过程建模模型互补性,以覆盖率最大化为目标,自适应选择模型子集,在RouterEval上实现更高覆盖率和更低冗余,同时保持灵活推理成本。
中文摘要 AI 辅助
现有的大语言模型(LLM)路由方法独立地对LLM进行评分,以选择前k个模型。然而,这忽略了模型之间的相关性,并强制施加了刚性的计算预算。因此,路由器常常选择共享失败模式的冗余模型,限制了整体的成功概率。为了解决这个问题,我们提出了FlexRouter,一个显式建模模型互补性的路由框架。FlexRouter优化“答案覆盖率”,最大化至少一个被选模型产生正确响应的概率。这一目标与实际的推理流程一致,在该流程中生成多个候选输出,并由下游验证器或用户选择最终输出。我们将路由问题表述为面向覆盖率的子集选择问题,并使用行列式点过程(DPPs)对路由策略进行建模,DPPs自然能够捕捉模型的胜任力和冗余性。为了直接优化覆盖率而无需真实目标子集,我们引入了一种基于失败集边缘化的训练目标。在推理过程中,我们采用基于边缘对数行列式增益的贪心策略,使路由器能够自适应地确定子集大小,而无需预设预算。在大型RouterEval基准上的大量实验表明,我们提出的FlexRouter在域内和域外任务上均比强基线实现了更高的覆盖率和更低的冗余度,同时保持了灵活的推理成本。
英文摘要
Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address this, we propose FlexRouter, a routing framework that explicitly models model complementarity. FlexRouter optimizes for \textit{answer coverage}, maximizing the probability that at least one selected model yields a correct response. This objective aligns with practical inference pipelines where multiple candidate outputs are generated and a downstream verifier or user selects the final one. We formulate routing as a coverage-oriented subset selection problem and model the routing policy using Determinantal Point Processes (DPPs), which naturally capture both model competence and redundancy. To directly optimize coverage without requiring a ground-truth target subset, we introduce a training objective based on marginalizing over failure sets. During inference, we employ a greedy strategy based on marginal log-determinant gains, enabling the router to adaptively determine subset sizes without a predefined budget. Extensive experiments on the large-scale RouterEval benchmark demonstrate that our proposed FlexRouter achieves higher coverage with lower redundancy across both in-domain and out-of-domain tasks than strong baselines while maintaining flexible inference cost.
发表机构
- Virginia Tech(弗吉尼亚理工大学)
- University of Southern California(南加州大学)
- Adobe Research(Adobe研究院)
- Dolby Labs(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。