arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TokenRouter:面向 token 级大语言模型路由的高效服务系统

TokenRouter: Efficient Serving System for Token-Level LLM Routing

Tianyu Fu, Tengxuan Liu, Ruoxi Wang, Yixin Dong, Yi Ge, Yichen You, Yu Wang

arXiv 2610.12242首次发表:更新:

发表机构

Tsinghua University; Carnegie Mellon University(清华大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TokenRouter 是面向 token 级 LLM 路由的高效服务系统,采用请求中心编程与模型中心执行的架构,通过延迟批处理调度器提升解码吞吐量,较现有系统实现 2.01-64.15 倍的效率提升。

AI 中文摘要

大语言模型(LLM)路由将推理工作分配至不同模型,推进了 LLM 服务的成本-质量帕累托前沿。尽管会话或查询级的粗粒度路由已在生产系统中广泛应用,但近期算法研究表明,细粒度的 token 级路由可带来显著的效率与质量提升。然而,现有系统难以高效支撑 token 级路由推理,其基于单 LLM 的假设会导致 token 级路由下出现严重的步骤失同步、频繁的批处理准入延迟,还会给开发者带来极高的实现复杂度。为应对这些挑战,我们设计了 TokenRouter——一个高效且开发者友好的 token 级 LLM 推理服务系统。TokenRouter 遵循「以请求为中心编程、以模型为中心执行」的原则:开发者从单个请求的视角描述路由逻辑,运行时则为每个 LLM 启动一个子服务器并异步分发请求。每个子服务器采用延迟批处理调度器,其最优超参数由系统的数学吞吐量模型推导得出。在各类路由算法、工作负载及模型组合下,TokenRouter 的解码吞吐量较现有系统提升了 2.01-64.15 倍,大幅推进了 token 级 LLM 路由的服务效率。我们的代码可在该 https URL 获取。

英文摘要

Large language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithmic work shows that fine-grained token-level routing can yield substantial efficiency and quality gains. However, efficiently serving token-level routed inference poses significant challenges to existing systems. Built on single-LLM assumptions, current systems suffer from severe step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address these challenges, we design TokenRouter, an efficient and developer-friendly serving system for token-level routed LLM inference. TokenRouter follows the principle of request-centric programming, model-centric execution: developers describe routing logic from the perspective of a single request, while the runtime launches a subserver for each LLM and dispatches requests asynchronously. Each subserver employs a delayed-batching scheduler, whose optimal hyperparameters are derived from a mathematical throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 2.01-64.15x higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing. Our code is available at https://github.com/thu-nics/TokenRouter.

CommentsAccepted by NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑