AI 中文总结
研究硬KV缓存内存预算下批量LLM推理的非先知式调度,提出基于状态感知路由框架的常数竞争算法,能处理任意提示和响应长度,对完工时间和总完成时间有常数竞争保证。
AI 中文摘要
我们研究了在硬键值(KV)缓存内存预算下,用于批量大语言模型(LLM)推理的非先知式调度。每个请求有已知的提示长度但未知响应长度,其内存占用包括固定提示部分和随解码令牌增长的响应部分。调度器在每个解码轮选择可行的活动请求批次,驱逐请求会浪费之前的计算。目标是针对知道所有响应长度的最优先知式调度最小化总完成时间。我们提出了首个针对任意提示长度和响应长度且无额外假设的常数竞争算法。该算法基于新颖的状态感知路由框架,由专门子调度器处理不同内存增长几何形状,元调度器跨它们分时内存预算并动态路由每个作业。此框架还为完工时间和在线到达下的总完成时间提供常数竞争保证。
英文摘要
We study non-clairvoyant scheduling for batched Large Language Model (LLM) inference under a hard Key-Value (KV) cache memory budget. Each request has a known prompt length but an unknown response length, and its memory footprint comprises a fixed prompt component together with a response component that grows with each decoded token. At each decoding round, the scheduler chooses a feasible batch of active requests; evicting a request discards its accumulated cache states, wasting prior computation. The goal is to minimize total completion time against the optimal clairvoyant schedule that knows all response lengths. We present the first constant-competitive algorithm for arbitrary prompt lengths and arbitrary response lengths with no additional assumptions. Rather than relying on a single universal scheduling policy, our algorithm is built on a novel regime-aware routing framework. Specialized sub-schedulers handle different memory-growth geometries, while a meta-scheduler time-shares the memory budget across them and dynamically routes each job as its execution progressively reveals its behavior. This framework also yields constant-competitive guarantees for makespan and for total completion time under online arrivals.