AI 中文总结
介绍用于大语言模型服务的轻量级框架PEEK,核心方法包括维护基数树、双路匹配等,贡献是在缓存命中率、处理帧定时等方面相比各引擎基线有显著提升,且在无可用前缀结构的工作负载上与基线相当。
AI 中文摘要
我们提出PEEK,一种用于在线(流式)和离线(批处理)大语言模型服务的轻量级调度和逐出框架;本文聚焦在线模式。PEEK在待处理队列上维护增量基数树,暴露现有引擎未呈现的前缀共享集群。低开销双路匹配将树与引擎前缀缓存匹配,为每个等待请求产生最长前缀匹配……
英文摘要
We present PEEK, a lightweight scheduling and eviction framework for both online (streaming) and offline (batch) LLM serving; this paper focuses on the online regime. PEEK maintains an incremental radix tree over the pending queue, exposing prefix-sharing clusters no existing engine surfaces. A low-overhead dual-walk matches the tree against the engine's prefix cache to yield longest-prefix-match for every waiting request; PEEK then admits cluster pioneers first so siblings inherit the freshly cached prefix, a co-designed eviction hook protects blocks ancestral to queued demand, and a multi-lane stride scheduler bounds starvation. On SGLang and vLLM across five workloads up to 4$\times$H100 (DP=2 over TP=2), PEEK delivers up to 3.0$\times$/2.6$\times$ cache hit, 7.9$\times$/7.1$\times$ TTFT, 6.7$\times$/5.5$\times$ E2E, and 3.6$\times$/4.5$\times$ throughput gains over each engine's strongest stock baseline (SGLang/vLLM), while matching baselines within noise on workloads with no exploitable prefix structure. Wins hold as KV-cache pressure and inference parallelism scale.
Comments26 pages, 21 figures, 21 tables. Preprint, under review