KV 缓存工作集:LLM 推理系统的在线容量规划
The KV Cache Working Set: Online Capacity Planning for LLM Inference Systems
- Kingsoft Cloud(金山云)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM服务中前缀缓存容量规划难题,提出在线分析器KVSET,利用Mattson栈算法高效估计KV缓存工作集,确定最小容量,经生产轨迹验证准确,支持在线与离线分析。
AI中文摘要:
前缀缓存对于高效的大型语言模型(LLM)服务至关重要,尤其是对于智能体工作负载,这类负载会随着对话和工具使用历史的增长而反复调用模型。通过重用先前处理过的前缀的键值(KV)状态,前缀缓存避免了冗余的预填充计算。然而,其有效性取决于保留足够大的 KV 缓存状态集合。为了保留所有历史 KV 状态而配置足够的缓存,成本过高且往往不必要,而容量不足则会显著降低缓存命中率。因此,确定 KV 缓存工作集(即实现目标命中率所需的最小缓存容量)对于高效的缓存配置和系统设计至关重要。我们提出了 KVSET,一种在线分析器,用于估计 LLM 服务工作负载的 KV 缓存工作集。KVSET 使用 Mattson 栈算法,在广泛的缓存容量范围内高效估计缓存命中率。对于每个 KV 缓存页,KVSET 计算其 LRU 栈距离,并将其与每个候选容量的页号进行比较。这种比较决定了该页在每个容量下是否会命中,而无需独立模拟每种容量配置。因此,KVSET 大幅降低了传统逐容量模拟的计算和内存开销,使在线工作集分析变得实用。KVSET 还根据实现目标命中率所需的前缀页中的最大 LRU 深度来确定最小缓存容量。我们使用从生产 LLM 工作负载中收集的轨迹验证了 KVSET,并表明其估计值与真实缓存部署的测量结果非常接近。开源实现支持在线请求处理和离线轨迹重放。
英文摘要:
Prefix caching is critical for efficient large language model (LLM) serving, particularly for agentic workloads that repeatedly invoke the model with a growing conversation and tool-use history. By reusing the key-value (KV) states of previously processed prefixes, prefix caching avoids redundant prefill computation. Its effectiveness, however, depends on retaining a sufficiently large set of KV cache states. Provisioning enough cache to preserve all historical KV states is prohibitively expensive and often unnecessary, whereas insufficient capacity can substantially degrade the cache hit rate. Determining the KV cache working set, defined as the minimum cache capacity required to achieve a target hit rate, is therefore essential for efficient cache provisioning and system design. We present KVSET, an online analyzer that estimates the KV cache working set of LLM serving workloads. KVSET uses the Mattson stack algorithm to efficiently estimate cache hit rates across a wide range of cache capacities. For each KV cache page, KVSET computes its LRU stack distance and compares it with the page number of each candidate capacity. This comparison determines whether the page would be a hit at each capacity without independently simulating every capacity configuration. KVSET therefore substantially reduces the computational and memory overhead of conventional capacity-by-capacity simulation and makes online working-set analysis practical. KVSET further determines the minimum cache capacity based on the maximum LRU depth among the prefix pages required to achieve the target hit rate. We validate KVSET using traces collected from production LLM workloads and show that its estimates closely match measurements from real cache deployments. The open-source implementation supports both online request processing and offline trace replay.