LLM Inference Under Bursty Workload Distribution: Modifying the WAIT Algorithm
突发工作负载分布下的大语言模型推理:对WAIT算法的改进
机构 * Indian Institute of Technology Dharwad(印度达尔瓦德印度理工学院) ; Indian Institute of Technology Kanpur(印度坎普尔印度理工学院)
专题命中 效率与部署 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.LG
AI总结 该研究针对LLM推理中突发工作负载的问题,改进了WAIT算法,通过在线估计请求强度提升了低到达率变化场景下的吞吐量,且延迟与现有基准相当。