物理分区的KVCache格式:MoE推理中CPU-GPU负载均衡
Physically Partitioned KVCache Format for CPU--GPU Load Balancing in MoE Inference
浏览论文内容
中文总结 AI 辅助
针对MoE长上下文推理中KVCache溢出到CPU的负载均衡问题,提出InplaceKVCache抽象与WriteScope调度,通过写入时固定物理驻留实现无数据移动的动态均衡,在多个模型上取得1.4-2.5倍加速。
中文摘要 AI 辅助
使用混合专家(MoE)模型的单GPU长上下文推理需要将键值缓存(KVCache)溢出到CPU内存。溢出的KV服务于两个互补的目的——传输到GPU进行注意力计算,或在CPU上就地计算——这要求相反的物理状态。两者之间的最优划分随工作负载而变化,然而现有的KVCache抽象仅提供对单一物理状态的单体对象的存储语义,无法表达动态负载均衡。我们提出InplaceKVCache,这是第一个KVCache抽象,其格式在写入时固定每个字节的物理驻留位置,从而可以在放置后无需移动数据即可调整CPU-GPU负载均衡。它通过沿两个维度——设备亲和性和访问模式——的四区域布局实现这一点,将负载均衡转化为纯调度。基于此抽象,WriteScope沿序列维度划分CPU-GPU份额,一个可移植的roofline性能模型确定随序列长度演变的最优CPU份额,并通过在线反馈跟踪CPU成本漂移。在三个MoE模型(DeepSeek-V2-Lite、Qwen3-30B-A3B、Mixtral-8×7B)上,在32 GB VRAM预算下,WriteScope支持1M token聚合规模下的端到端推理。在长上下文区域(≥8K)中,它在A100上实现了1.5倍至2.5倍的几何平均加速,在V100上实现了1.4倍至1.7倍的几何平均加速,相比四个复现的基线,而vLLM、SGLang和KTransformers即使KV预算加倍也失败。DeepSeek-V4-Flash案例研究验证了与原生稀疏注意力的组合。
英文摘要
Single-GPU long-context inference with Mixture-of-Experts (MoE) models requires spilling the key-value cache (KVCache) to CPU memory. The spilled KV serves two complementary purposes---transferring to the GPU for attention computation, or computing in-place on the CPU---which demand opposing physical states. The optimal split between them varies with workload, yet existing KVCache abstractions offer only storage semantics over a monolithic object of a single physical state, and cannot express dynamic load balancing. We propose InplaceKVCache, the first KVCache abstraction whose format fixes each byte's physical residency at write time, so that the CPU--GPU load balance can be adjusted without moving data after placement. It realizes this as a four-region layout along two dimensions---device affinity and access pattern---turning load balancing into pure scheduling. Built on this abstraction, WriteScope splits CPU--GPU shares along the sequence dimension, and a portable roofline performance model determines the optimal CPU share as sequence length evolves, with online feedback tracking CPU cost drift. On three MoE models (DeepSeek-V2-Lite, Qwen3-30B-A3B, Mixtral-8$\times$7B) with a 32~GB VRAM budget, WriteScope supports end-to-end inference at the 1M-token aggregate scale. In the long-context regime ($\ge$8K), it achieves geometric-mean speedups of $1.5\times$--$2.5\times$ on A100 and $1.4\times$--$1.7\times$ on V100 over four reproduced baselines, while vLLM, SGLang, and KTransformers fail even with a doubled KV budget. A DeepSeek-V4-Flash case study validates composition with native sparse attention.
发表机构
- National University of Defense Technology(国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。