arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09225cs.CRcs.AI

管控KV缓存:防止多租户大语言模型推理中的时序侧信道泄露

Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference

Tejasvi C. Addagada

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对多租户LLM推理中KV缓存引发的时序侧信道泄露问题,提出KVGov治理层,结合盐值机制与ORIGAMI调度器,可抵御三类攻击,同时保留93%的缓存效率,在真实硬件与模拟环境中验证了防御效果。

中文摘要 AI 辅助

键值(KV)缓存是现代大语言模型(LLM)推理中主要的吞吐量优化手段,支持跨请求的前缀复用。在多租户部署场景下,该缓存由所有租户共享,由此产生了时序侧信道:恶意租户可通过探测缓存命中时延重构其他租户的私有提示词。目前已有三种公开的攻击方法——PROMPTPEEK、EarlyBird和InputSnatch,针对未受保护的vLLM和SGLang,攻击成功率最高可达100%,具体比例随缓存架构和提示词结构变化。本文提出KVGov,一种治理层,通过单一机制覆盖所有三类攻击的前缀缓存路径。该方案为每个主体生成盐值σ_p = HMAC_K(密钥, 主体ID),作为块哈希链的种子,使不同主体的缓存密钥在密码学层面完全不重叠。一项消融实验(N=1000次试验,随机种子2026,确定性判定)证实该盐值是必要且充分的组件。KVGov还引入ORIGAMI,一种Stackelberg水填充审计调度器,在实际租户异质性(基尼系数0.63)下将攻击者的预期效用降低12.6%;同时通过进化稳定性分析得出,当攻击者流行率低于31.6%的临界点时,全局缓存仍可保持稳定。在真实硬件(Qwen2.5-7B-Instruct、vLLM 0.26.0、NVIDIA A100)上,我们测得经门控验证的冷启动/缓存首包生成时间(TTFT)比值为0.22,证实该信道在生产规模下可被利用;防御方案本身则在基于该测量值校准的模拟环境中进行评估。我们在独立技术栈(Apple Metal上的相关实现,比值为0.093)上复现了该信道。最后,隔离性与缓存效率无需兼顾冲突:相关信息仅存在于提示词发生分歧的位置,因此在该边界处注入盐值而非哈希链根,可保留约93%的前缀缓存收益,且不会产生跨主体信号。

英文摘要

The key-value (KV) cache is the primary throughput optimization in modern large language model (LLM) inference, enabling prefix reuse across requests. In multi-tenant deployments this cache is shared across tenants, creating a timing side channel: an adversarial tenant can reconstruct another tenant's private prompt by probing cache-hit latency. Three published attacks exploit it -- PROMPTPEEK, EarlyBird and InputSnatch -- reaching up to 100% attack success rate against unprotected vLLM and SGLang, with rates varying by cache architecture and prompt structure. We present KVGov, a governance layer addressing all three attack families' prefix-cache paths under one mechanism. A per-principal salt sigma_p = HMAC_K(secret, principal_id) seeds the block-hash chain, making cache keys cryptographically disjoint across principals. An ablation (N=1000 trials, seed 2026, deterministic judges) isolates this salt as the necessary and sufficient component. KVGov adds ORIGAMI, a Stackelberg water-filling audit scheduler that reduces adversary expected utility by 12.6% at realistic tenant heterogeneity (Gini 0.63), and an evolutionary stability analysis giving a 31.6% adversary-prevalence tipping point below which global caching remains stable. On real hardware (Qwen2.5-7B-Instruct, vLLM 0.26.0, NVIDIA A100) we measure a gate-verified cold/cached TTFT ratio of 0.22, confirming the channel is exploitable at production scale; the defense itself is evaluated in simulation calibrated to those measurements. We replicate the channel on an independent stack (llama.cpp on Apple Metal, ratio 0.093). Finally, isolation and cache efficiency need not conflict: identifying information resides only where prompts diverge, so injecting the salt at that boundary rather than the chain root retains an estimated 93% of the prefix-cache benefit with no cross-principal signal.

补充信息

↑