arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06853cs.CRcs.AI

表征多租户LLM服务中KV缓存时序侧信道由争用引起的可靠性崩溃

Characterizing Contention-Induced Reliability Collapse in KV-Cache Timing Side Channels for Multi-Tenant LLM Serving

Rana Abu Bakar

AI总结:

该研究通过实验表征多租户LLM服务中KV缓存时序侧信道攻击的可靠性,发现争用导致可靠性崩溃,且安静系统测量会高估攻击可靠性。

AI中文摘要:

共享键值(KV)缓存复用可改善大语言模型(LLM)服务,但也可能产生一个时序侧信道,揭示某个前缀是否已被缓存。先前工作表明此类攻击可行,但在现实多租户争用下的可靠性尚不明确。我们通过在实时共享LLM服务系统上的七项实验研究该问题。在运行DeepSeek-R1-Distill-Llama-8B的NVIDIA GB10上的vLLM服务器中,平均Cohen's d从无合成工作负载时的0.7789降至有两个工作负载时的0.2109(t=8.412),而更高的工作负载数量未导致统计上可检测的进一步损失。一项120次运行的稀疏重叠实验将最佳断点置于测量范围边界(tau=0,95%置信区间[0.000,0.113]),支持环境负载与高负载状态之间的转变,而非内部物理阈值。AUROC从环境状态下的0.650降至约61%重叠时的0.531,并在饱和时部分恢复至0.574。并发深度方差是效应量(r=-0.416)和命中一致性(r=-0.637)的最强测量相关因素。交错控制保持了相同的非单调排序。主要崩溃也在真实的两节点、两GPU张量并行vLLM设置中复现,平均d从3.418降至0.511(p<0.01)。两项SGLang试点在统计上无定论。总体而言,KV缓存时序可靠性强烈依赖于负载状态和服务堆栈,在安静系统上的测量可能高估实际攻击可靠性。

英文摘要:

Shared key--value (KV) cache reuse improves large language model (LLM) serving, but it can also create a timing side channel that reveals whether a prefix is already cached. Previous work shows that such attacks are possible, but their reliability under realistic multi-tenant contention is less understood. We study this problem through seven experiments on live shared LLM-serving systems. On a vLLM server running DeepSeek-R1-Distill-Llama-8B on NVIDIA GB10, mean Cohen's d drops from 0.7789 with no synthetic workers to 0.2109 with two workers (t=8.412), while higher worker counts cause no statistically detectable further loss. A 120-run sparse-overlap experiment places the best breakpoint at the boundary of the measured range (tau=0, 95% CI [0.000,0.113]), supporting an ambient-versus-loaded regime change rather than an internal physical threshold. AUROC falls from 0.650 at ambient to 0.531 near 61% overlap and partially recovers to 0.574 at saturation. Concurrency-depth variance is the strongest measured correlate of effect size (r=-0.416) and hit consistency (r=-0.637). An interleaved control preserves the same non-monotonic ordering. The main collapse is also reproduced on a real two-node, two-GPU tensor-parallel vLLM setup, where mean d falls from 3.418 to 0.511 (p<0.01). Two SGLang pilots are statistically inconclusive. Overall, KV-cache timing reliability depends strongly on the load regime and serving stack, and measurements on quiet systems can overestimate operational attack reliability.

↑