发表机构
Indian Institute of Technology Dharwad; Indian Institute of Technology Kanpur(印度达尔瓦德印度理工学院; 印度坎普尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM推理中突发工作负载的问题,改进了WAIT算法,通过在线估计请求强度提升了低到达率变化场景下的吞吐量,且延迟与现有基准相当。
AI 中文摘要
ChatGPT、Claude等大语言模型(LLM)被广泛用于信息检索和问题解决。近期研究聚焦于改进调度算法以提升吞吐量并保持低延迟,但这些方法通常假设请求到达遵循泊松过程且速率恒定,该假设无法反映现实流量固有的突发性和动态性。我们对最先进的WAIT算法[1]提出轻量级扩展,无需先验流量知识即可适应随时间变化的到达率。该算法基于观测到的到达间隔时间对请求强度进行在线估计。我们采用基于马尔可夫调制泊松过程(MMPP)的合成工作负载(含多种请求类型)进行仿真评估,结果表明,在评估的低到达率变化场景中,所提方法实现了比Sarathi-Serve[2]、ORCA[3]和vLLM[4]更高的吞吐量,同时保持了相当的延迟。
英文摘要
Large Language Models (LLMs) such as ChatGPT and Claude are widely used for information retrieval and problem-solving. Recent work has focused on improving scheduling algorithms to boost throughput while maintaining low latency. However, these approaches often assume Poisson request arrivals with constant rates - an assumption that fails to reflect the inherently bursty and dynamic nature of real-world traffic. We propose a lightweight extension to the state-of-the-art WAIT algorithm [1], which adapts to time-varying arrival rates without prior traffic knowledge. The proposed algorithm performs online estimation of request intensity based on observed interarrival times. Using Markov Modulated Poisson Process (MMPP)-based synthetic workloads with diverse request types, we conduct a simulation-based evaluation demonstrating that the proposed method achieves higher throughput than Sarathi-Serve [2], ORCA [3], and vLLM [4] in the evaluated low arrival-rate shift scenarios while maintaining comparable latency.