发表机构
School of Artificial Intelligence, Nanjing University; National Key Laboratory for Novel Software Technology, Nanjing University; The Hong Kong University of Science and Technology; SinapisAI(南京大学人工智能学院; 南京大学计算机软件新技术全国重点实验室; 香港科技大学; SinapisAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM路由中监督成本高昂的问题,提出稀疏监督框架SaveRouter,选择性获取反馈并共享能力信息,在四个基准上仅用33%-41%训练反馈保持路由质量,并将盈亏平衡部署量降低1.9-9.5倍。
AI 中文摘要
大语言模型(LLM)路由通过将每个查询分配给合适的模型来降低服务成本,同时保持响应质量。然而,学习这样一个路由器通常需要在历史查询上执行多个候选模型以收集查询-模型质量反馈,从而在部署前产生不可忽视的监督成本。现有工作主要关注服务时效率,忽略了由此产生的节省是否足以回收这笔前期支出。我们进一步观察到,路由质量往往在收集完所有查询-模型反馈之前就已饱和,这表明密集监督在经济上可能过度供给。我们提出SaveRouter,一种稀疏监督路由框架,它选择性地获取信息丰富的模型反馈,并在相关查询之间共享能力信息,同时保留查询级细化以实现细粒度路由。我们通过联合考虑监督支出和后续服务节省来评估路由。在四个路由基准上,主要设置仅使用约33%至41%的可用训练反馈,同时保持竞争性或更好的路由质量,并将盈亏平衡部署量相比最快的传统路由器减少了约1.9至9.5倍。进一步分析表明,获取更多监督并不总是在经济上更优:使服务成本最小化的监督水平可能与实现最早回本的监督水平不同。我们的代码可在以下网址公开获取:此https URL。
英文摘要
Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query--model feedback is collected, suggesting that dense supervision can be economically over-provisioned. We propose SaveRouter, a sparse-supervision routing framework that selectively acquires informative model feedback and shares capability information across related queries, while retaining query-level refinement for fine-grained routing. We evaluate routing by jointly accounting for supervision expenditure and subsequent serving-time savings. Across four routing benchmarks, the main setting uses only about 33--41% of available training feedback while maintaining competitive or better routing quality, and reduces the break-even deployment volume by approximately 1.9--9.5 times compared with the fastest conventional router. Further analysis shows that acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that achieves the earliest payback. Our code is publicly available at https://github.com/LAMDA-Model-Reuse/SaveRouter.