共享KV缓存用于复制的27B推理:正确性失败与性能边界
Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries
- UNSW Sydney(悉尼新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过工程案例验证共享KV缓存的正确性修复,并测量27B推理副本间延迟,发现交替放置可显著降低首令牌时间,而固定放置收益甚微。
AI中文摘要:
共享主机内存缓存可以避免请求在推理副本之间移动时的重复预填充。其有效性既取决于正确的状态传输,也取决于丢失的前缀局部性。我们研究了两个单GPU 27B vLLM副本共享一个256 GiB LMCache池。在采用现有的打包页补丁后,我们隔离了一个原始指针回退,该回退忽略了对当前CUDA流的依赖。受控字节测试在施加延迟下失败,而在恢复依赖时通过;现有的混合分配器提供了一条可行的部署路径。全池分配检查和服务回归完成了验证。一个四块OFF-ON-ON-OFF比较包含两个块对内的768个测量请求。跨副本的首个内容令牌中位时间在128k输入时从31.715秒降至0.605秒,在256k时从92.047秒降至0.790秒。交替副本的六轮合成会话在32k和128k的初始上下文下分别改善了约35%和45%,而固定放置则几乎没有益处。这项工程案例研究确定了实用的验证步骤以及共享缓存产生效益的局部性条件。
英文摘要:
Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality. We study two single-GPU 27B vLLM replicas sharing a 256 GiB LMCache pool. After adopting an existing packed-page patch, we isolate a raw-pointer fallback that omits the dependency on the current CUDA stream. Controlled byte tests fail under an imposed delay and pass when the dependency is restored; the existing mixed allocator provides a working deployment path. Full-pool allocation checks and service regression complete the validation. A four-block OFF-ON-ON-OFF comparison contains 768 measured requests within two block pairs. Median cross-replica time to first content token falls from 31.715 to 0.605 seconds at 128k input and from 92.047 to 0.790 seconds at 256k. Six-turn synthetic sessions alternating replicas improve by approximately 35% and 45% at initial contexts of 32k and 128k, while fixed placement shows little benefit. This engineering case study identifies practical validation steps and the locality conditions in which shared caching pays off.