arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SMetric:以平衡的会话中心调度重新思考面向服务代理的大语言模型调度

SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling

Jiahao Wang, Kaizhan Lin, Kaixi Zhang, Jinbo Han, Xingda Wei, Sijie Shen, Chenguang Fang, Wenyuan Yu, Rong Chen, Haibo Chen

arXiv 2607.08565首次发表:更新:

发表机构

Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University; Alibaba Group(并行与分布式系统研究所,上海交通大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究代理服务中LLM调度问题,提出基于全局层KV存储和会话内局部性的见解,设计SMetric平衡会话中心调度,相比先进调度器提升集群TPS及预填充TPS,降低每个令牌延迟。

AI 中文摘要

大语言模型(LLM)调度对服务至关重要,但现有设计在面向代理服务(由代理而非人类发出LLM请求)方面的适配情况尚不明晰。这在两方面改变了工作负载:代理仅依据完整响应行动,使集群每秒令牌数(TPS)成为主要目标,放宽而非消除了每个令牌的延迟要求;请求共享大量键值(KV),在来自百度的生产跟踪中,KV重用超过请求令牌的80%,而在聊天中为54 - 62%。本文首先对基于两个真实世界跟踪的代理请求调度进行了系统研究。发现为增加KV重用,现有调度器过度优先将请求路由到缓存其KV的实例,导致一些实例过载而其他闲置,限制了TPS。因此提出两个关键见解:借助全局层KV存储,负载平衡无需牺牲所有KV重用;利用工作负载的会话内局部性,平衡一小部分请求(每个代理会话的第一个请求)就足以平衡集群,而不牺牲本地实例上的大部分KV重用。SMetric通过平衡的会话中心调度实现了这些见解:它纯粹为了负载平衡路由每个会话的第一个请求,并以缓存感知方式路由后续请求,在保持全局层需求较低的同时保持负载平衡和本地重用。使用会话轮次信息作为调度指标是有意为之:它仅从用户输入高效准确地得出,使调度器保持简洁无状态。在与全局存储的预填充 - 解码共置情况下,SMetric比最先进的调度器将集群TPS提高了10 - 16%,在分解情况下将预填充TPS提高了2 - 34%,并且每个令牌的延迟也更好。

英文摘要

LLM scheduling is critical to serving, yet how well existing designs fit agentic serving--where agents, not humans, issue the requests--remains unclear. Agents shift the workload in two ways: they consume many more tokens than humans, so the cluster must provide high throughput (TPS) at low latency; and their requests reuse far more KV\$ than chat. Existing schedulers still trade off load balance against KV\$ reuse: cache-aware schedulers may crowd requests onto the few instances caching the KV\$, leaving the rest idle, while balanced schedulers may lose the opportunity for reuse, which is costly at a high reuse ratio. We thus present two key insights: (1) with a global-tier KV\$ store, pursuing load balance need not compromise KV\$ reuse, though the slower global tier must be used with care; and (2) given the agent's intra-session locality, routing requests by their sessions can balance the load with high KV\$ reuse. A key challenge in realizing session-centric scheduling is that the scheduler must identify a request's session statelessly, which is difficult for model providers serving arbitrary agents. SMetric addresses this with differential scheduling based on two indicators derived from the request itself, the session turn and the local KV\$ hit: it schedules first-turn requests for load balance, and sticks follow-ups to the instance with the highest local hit for high KV\$ reuse. As sessions differ widely in size, SMetric sticks a follow-up only if the instance can serve it within its SLO, and otherwise migrates the session to the least-loaded instance to prevent many long sessions from eventually imbalancing the load. Evaluated on real-world traces, SMetric improves the peak TPS by 9-15% under prefill-decode colocation with a provisioned global tier and the peak prefill TPS by 9% under disaggregation over state-of-the-art schedulers, also with lower latency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑