面向多轮 MoE 服务的动态 HBM 重分区
Dynamic HBM Repartitioning for Multi-Turn MoE Serving
浏览论文内容
中文总结 AI 辅助
VAMP 框架通过运行时动态调整 MoE 模型权重与 KV 缓存的 HBM 边界,在 KV 分配不足时选择最优策略,显著降低多轮服务延迟并提升吞吐量。
中文摘要 AI 辅助
长时间运行的多轮请求会累积可复用的键值(KV)状态。一旦该状态超过固定的 GPU KV 缓存分配,服务系统就会驱逐可复用的前缀,重复预填充工作,并可能抢占请求。这种压力对于混合专家(MoE)模型尤为严重:其专家权重占据了大部分 GPU 高带宽内存(HBM),尽管每个令牌仅激活稀疏的专家子集。权重与 KV 缓存之间的静态边界阻止了服务系统在对话增长时利用专家内存来保留可复用状态。我们提出了 VAMP,一个 MoE 服务框架,它在运行时改变这一 HBM 边界。当 KV 分配无法满足时,VAMP 比较三种替代方案的估计未来工作量:从主机内存暂存专家权重、驱逐可能需要重新预填充的缓存前缀、或抢占并重新调度请求。然后,当该动作具有最低估计惩罚时,它将一个有界的专家权重区域转换为 KV 缓存容量。CUDA 虚拟内存管理页面重映射在不复制驻留 KV 数据的情况下执行此转换。我们在 vLLM 服务引擎中实现了 VAMP,并在记录和受控的多轮工作负载上评估了 Qwen3-Next-80B。在对记录的 2,103 轮 SWE-bench 智能体工作负载的五次重放中,VAMP 在最大专家卸载比例为 15% 的情况下,将首令牌时间(TTFT)的 p90 从 26.1 秒降低到 1.10 秒(降低 23.6 倍),并将请求吞吐量相对于未修改的 vLLM 提高了 20.7%,同时每输出令牌时间(TPOT)增加了 31.1%。
英文摘要
Long-running multi-turn requests accumulate reusable key-value (KV) state. Once this state exceeds a fixed GPU KV-cache allocation, serving systems evict reusable prefixes, repeat prefill work, and may preempt requests. This pressure is particularly acute for Mixture-of-Experts (MoE) models: their expert weights occupy most GPU high-bandwidth memory (HBM), even though each token activates only a sparse subset of experts. A static boundary between weights and the KV cache prevents serving systems from using expert memory to preserve reusable state as conversations grow. We present VAMP, an MoE serving framework that changes this HBM boundary at runtime. When a KV allocation cannot be satisfied, VAMP compares the estimated future work of three alternatives: staging expert weights from host memory, evicting cached prefixes that may require re-prefill, or preempting and rescheduling requests. It then converts a bounded expert-weight region into KV-cache capacity when that action has the lowest estimated penalty. CUDA Virtual Memory Management page remapping performs this conversion without copying resident KV data. We implement VAMP in the vLLM serving engine and evaluate Qwen3-Next-80B on recorded and controlled multi-turn workloads. Across five replays of a recorded 2,103-turn SWE-bench agent workload, VAMP with a 15% maximum expert-offloading ratio reduces time-to-first-token (TTFT) p90 from 26.1 s to 1.10 s (23.6 times) and increases request throughput by 20.7% relative to unmodified vLLM, while increasing time per output token (TPOT) by 31.1%.
发表机构
- UC San Diego(加州大学圣地亚哥分校)
- SK hynix(SK海力士)
机构由 AI 辅助整理,请以论文原文为准。