arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15359cs.AI

MAPS:面向大语言模型服务的内存感知预测调度框架

MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving

  • College of Intelligence and Computing, Tianjin University(天津大学智能与计算学部)
  • The International Joint Institute of Tianjin University, Tianjin University(天津大学国际联合学院)
  • Faculty of Digital Economics and Managements, Tianjin University of Finance and Economics(天津财经大学数字经济与管理学院)

机构由 AI 辅助整理,请以论文原文为准。

Tiancheng Zhang, Yulin Chen, Yunfeng Zhao, Shaoyuan Huang, Cheng Zhang, Xiaofei Wang

AI总结:

针对分离式LLM服务中解码实例负载不均的问题,提出内存感知预测调度框架MAPS,通过设备辅助输出长度预测与不确定性校准实现安全调度,显著降低端到端延迟。

AI中文摘要:

个人设备上大语言模型(LLM)应用的激增给云服务基础设施带来了大量突发性工作负载。虽然预填充-解码分离(prefill-decode disaggregation)提高了吞吐量和可扩展性,但受内存限制的解码实例常常遭受持续的负载不平衡,因为请求到达云端时输出长度是未知的。为解决这一问题,我们提出了MAPS,一种专为分离式LLM服务量身定制的内存感知预测调度框架。MAPS执行设备辅助的推测性输出长度预测,并与云端预填充重叠进行,仅产生可忽略的延迟开销。为处理生成不确定性,MAPS应用不确定性感知校准来推导具有目标覆盖率的输出长度上界,从而实现安全的调度决策。基于这些上界,MAPS采用分层全局-局部调度策略,以缓解解码器间队列堆积和解码器内队头阻塞。在两个真实工作负载和两个LLM上的大量实验表明,MAPS显著优于三个最先进的系统,将平均端到端延迟降低42.6%,尾部延迟降低高达84.8%。

英文摘要:

The surge of large language model (LLM) applications on personal devices imposes massive, bursty workloads on cloud serving infrastructure. While prefill-decode disaggregation improves throughput and scalability, memory-bound decode instances often suffer from persistent load imbalance, as output lengths are unknown when requests arrive at the cloud. To address this, we propose MAPS, a Memory-Aware Predictive Scheduling framework tailored for disaggregated LLM serving. MAPS performs device-assisted speculative output length prediction overlapped with cloud-side prefilling, incurring negligible latency overhead. To handle generation uncertainty, MAPS applies uncertainty-aware calibration to derive output-length upper bounds with target coverage, enabling safe scheduling decisions. Building on these bounds, MAPS employs a hierarchical global-local scheduling strategy to mitigate inter-decoder queue buildup and intra-decoder head-of-line blocking. Extensive experiments on two real-world workloads and two LLMs show that MAPS significantly outperforms three state-of-the-art systems, reducing average end-to-end latency by 42.6 and tail latency by up to 84.8.

补充信息

↑