AI 中文总结
本文针对批量LLM服务中异构请求导致的资源分配低效问题,提出ISJL算法,实现高吞吐量的同时对齐最大驱动批次成本与token计量收入,达到3/4的竞争比下界,平衡了FCFS与LJF的优缺点。
AI 中文摘要
本文研究了批量大语言模型(LLM)服务中的资源分配低效问题:共享解码批次的异构请求会对彼此施加最大驱动的计算成本。由于批次步骤的挂钟成本主要由最大的活跃KV缓存 footprint 决定,与长请求同批次的短请求可能会经历与其自身 token 工作量不成比例的延迟和GPU资源消耗。我们将这一现象形式化为资源公平调度问题,开发了一种将批次内资源公平性与系统吞吐量关联起来的数学调度模型。所提出的公平性约束限制了同批次请求间解码进度(等价于KV缓存 footprint)的差异。基于该模型,我们设计了参数化混合批处理策略Insert-Short-Jobs-with-Limit(ISJL)算法,证明ISJL可达到3/4的全局竞争比下界。我们还在商业LLM API采用的token计量定价约定下,研究了资源公平调度的利润影响。数值实验表明,ISJL在FCFS(具有较大的批处理外部性)和LJF(成本对齐但牺牲批处理灵活性)之间占据有利的中间地带。因此,ISJL提供了一种双准则调度策略:它在保持高吞吐量的同时,使最大驱动的批次成本与token计量收入对齐。
英文摘要
This paper studies a resource-allocation inefficiency in batched large language model (LLM) serving: heterogeneous requests that share a decode batch impose max-driven computational costs on one another. Because the wall-clock cost of a batch step is largely governed by the largest active KV-cache footprint, a short request co-batched with a long request can experience latency and GPU-resource consumption disproportionate to its own token workload. We formalize this phenomenon as a resource-fair scheduling problem. We develop a mathematical scheduling model that connects within-batch resource fairness to system throughput. The proposed fairness constraint bounds the disparity in decode progress, equivalently KV-cache footprint, among co-batched requests. Based on this model, we design the Insert-Short-Jobs-with-Limit (ISJL) algorithm, a parameterized hybrid batching policy. We prove that ISJL achieves a global competitive-ratio lower bound of $3/4$. We further examine the profit implications of resource-fair scheduling under the token-metered pricing convention used by commercial LLM APIs. Numerical experiments show that ISJL occupies a favorable middle ground between FCFS, which has large batching externalities, and LJF, which is cost-aligned but sacrifices batching flexibility. Thus, ISJL provides a bi-criterion scheduling policy: it maintains high throughput while aligning max-driven batch cost with token-metered revenue.