发表机构
Griffith University; National University of Singapore; Macquarie University(格里菲斯大学; 新加坡国立大学; 麦考瑞大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究首次系统揭示LLM服务中共享KV缓存的来源盲复用问题,提出KV来源契约及绑定机制,在vLLM和SGLang中实现,消除不安全复用并保持低延迟。
AI 中文摘要
生产级LLM服务栈将推理引擎的本地前缀缓存与共享KV缓存层相结合,以实现跨集群的复用。本地缓存通过适配器、权重配置和共享域来区分请求,但共享层可能仅根据令牌内容和粗略的模型元数据来键控条目。这种边界抹除了来源信息,使得在不同计算或共享上下文下的相同令牌发生冲突。我们将这种组合缺口称为来源盲复用,并首次对其进行了系统性研究。对三个vLLM连接器的源代码审计确认了结构性遗漏,而运行时实验在vLLM和两个SGLang版本、来自7个家族的12个模型(0.5B-32B)以及超过160种配置中重现了该问题。跨适配器冲突将准确率从0.94降至0.64,不兼容的KV表示将推理准确率降至零,而盐值遗漏使得通过时序可识别93%的提示。我们将缺失的保证形式化为KV来源契约:对于声明的维度注册表,共享键必须在计算和共享来源上是单射的,且身份在各工作节点间保持稳定。因此,任何具有稳定身份的维度都可以添加,而无需连接器特定的键逻辑。一个规范描述符将每请求和每工作节点的来源绑定到查找和存储键中,而差分检查器检测改变KV状态但不改变键的维度。在vLLM和SGLang 0.5.20中跨三条缓存路径实现后,来源绑定消除了不安全复用,同时保留了合法共享。命中路径延迟变化保持在0.34毫秒以内且低于运行间差异;保留率随来源多样性增长。
英文摘要
Production LLM serving stacks combine an inference engine's local prefix cache with a shared KV-cache tier for fleet-wide reuse. The local cache distinguishes requests by adapter, weight configuration and sharing domain, but the shared tier may key entries only by token content and coarse model metadata. This boundary erases provenance and lets identical tokens under incompatible computational or sharing contexts collide. We call this composition gap provenance-blind reuse and present its first systematic study. A source audit of three vLLM connectors confirms the structural omission, while runtime experiments reproduce it across vLLM and two SGLang releases, 12 models from 7 families (0.5 B-32 B), and over 160 configurations. Cross-adapter collisions reduce accuracy from 0.94 to 0.64, incompatible KV representations reduce reasoning accuracy to zero, and salt omission enables 93% prompt identification from timing. We formalize the missing guarantee as the KV provenance contract: for a declared dimension registry, shared keys must be injective over computational and sharing provenance, with identities stable across workers. Any dimension with a stable identity can therefore be added without connector-specific key logic. A canonical descriptor binds per-request and per-worker provenance into lookup and store keys, while a differential checker detects dimensions that change KV state without changing the key. Implemented in vLLM and SGLang 0.5.20 across three cache paths, provenance binding eliminates unsafe reuse while preserving legitimate sharing. Hit-path latency changes remain within 0.34 ms and below run-to-run variation; retention grows with provenance diversity.
Comments14 pages, 3 figures