arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18253cs.AI

超越准确性和成本:针对动态工作负载的延迟感知大语言模型查询路由

Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads

发表机构卡内基梅隆大学 · 微软
查看机构详情
  • Carnegie Mellon University(卡内基梅隆大学)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

Shivam Patel, Akaash R. Parthasarathy, Ankur Mallick, Gauri Joshi

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对动态工作负载的大语言模型查询路由,设计轻量级延迟估计器,纳入延迟感知路由器,联合优化延迟、准确性和成本,实验表明该方法能在保持延迟的同时提升准确性 - 成本效用达40%。

中文摘要 AI 辅助

现代语言查询路由器通过为每个查询分配一个平衡响应质量和货币成本的模型来提高推理效率。然而,当前的查询路由器在很大程度上与延迟无关,没有考虑查询在模型实例处经历的生成延迟。在实践中,延迟通常由诸如轮询或加入最短队列等负载均衡策略控制,这些策略没有考虑模型准确性或推理成本。将查询延迟纳入路由具有挑战性,因为它不仅取决于查询的提示长度,还取决于模型实例当前的预填充和解码工作负载以及服务框架的调度和批处理策略。我们设计了一种轻量级延迟估计器,它在服务框架中模拟自回归令牌批处理,并估计查询的第一个令牌时间(TTFT)。我们将此延迟估计器纳入一个延迟感知路由器,该路由器在将查询分配给模型实例时联合优化延迟、准确性和成本。我们的实验结果表明,这种联合优化在保持与标准负载均衡方法相同延迟的同时,可使准确性 - 成本效用提高高达40%。

英文摘要

Modern language query routers improve inference efficiency by assigning each query to a model that balances response quality and monetary cost. However, current query routers are largely latency-agnostic and do not consider the generation latency experienced by queries at model instances. In practice, latency is often controlled by load-balancing policies such as round-robin or join-the-shortest-queue, which do not account for model accuracy or inference cost. Incorporating query latency into routing is challenging as it depends not only on the query's prompt length, but also on the current prefill and decode workload at the model instance and the scheduling and batching policy of the serving framework. We design a lightweight latency estimator that simulates autoregressive token batch processing in the serving framework and estimates the time-to-first-token (TTFT) of queries. We incorporate this latency estimator into a latency-aware router that jointly optimizes latency, accuracy, and cost when assigning queries to model instances. Our experimental results indicate that this joint optimization yields up to 40% improvement in accuracy--cost utility while maintaining the same latencies as standard load-balancing approaches.

↑