发表机构
Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出RouteFM,一种基于基础模型的LLM路由方法,通过行为上下文表征匿名模型并跨环境预训练,实现可迁移的路由能力,在MMR-Bench上以少量观测超越基线。
AI 中文摘要
大型语言模型(LLM)路由旨在将每个查询分配给异构候选池中最合适的模型,从而改善LLM推理的质量-效率权衡。现有的路由器通常通过局部拟合来学习:路由器针对特定的查询工作负载和候选池进行优化,并且随着路由环境的变化,通常需要额外的监督或重新训练。我们探讨是否可以从基础模型的角度来处理LLM路由,学习一种可复用的路由能力,该能力可跨任务、候选模型和部署条件进行泛化。为此,我们引入了RouteFM,它学习从行为上下文中表征匿名候选模型,并推断其针对特定目标的能力,而不是将路由决策绑定到固定的模型身份或单一环境。通过跨异构路由环境的 episodic 预训练,这种能力可以被冻结的路由器复用,并仅通过上下文适应新环境。实验证明了在领域、模态、候选池和上下文预算变化上的迁移,且在行为证据有限时收益最大。在未纳入预训练的MMR-Bench上,RouteFM在每位候选仅八次观测的情况下,比最强基线高出2.23个质量点。这些结果支持将LLM路由从重复的局部拟合转向“预训练一次,随处路由”的范式。我们的代码可在该 https URL 公开获取。
英文摘要
Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate pool, and often requires additional supervision or retraining as the routing environment changes. We ask whether LLM routing can instead be approached from a foundation-model perspective, learning a reusable routing capability that generalizes across tasks, candidate models, and deployment conditions. To this end, we introduce RouteFM, which learns to characterize anonymous candidate models from behavioral context and infer their target-specific capabilities, rather than binding routing decisions to fixed model identities or a single environment. Through episodic pretraining across heterogeneous routing environments, this capability can be reused by a frozen router and adapted to new environments through context alone. Experiments demonstrate transfer across changes in domains, modalities, candidate pools, and context budgets, with the largest gains when behavioral evidence is limited. On MMR-Bench, which is excluded from pretraining, RouteFM outperforms the strongest baseline by 2.23 quality points with only eight observations per candidate. These results support moving LLM routing from repeated local fitting toward a pretrain once, route anywhere paradigm. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/RouteFM.