Mira:使用自适应缓存和预测性专家分阶段加载的内存高效MoE推理
Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging
浏览论文内容
中文总结 AI 辅助
针对MoE模型在单GPU上推理时专家参数内存占用大、路由动态导致缓存效率低的问题,提出算法-系统协同设计Mira,通过预测性专家预取和定制量化缓存,实现5.71倍吞吐量提升。
中文摘要 AI 辅助
混合专家(MoE)模型是一种用于扩展模型容量的引人注目的架构,使其在资源受限的单GPU系统上部署时尤其具有吸引力。然而,这一优势难以实现,因为专家参数主导了内存使用,且令牌级路由是动态的、不可预测且偏斜的。先前使用卸载和缓存的工作从根本上是被动的,因为系统在移动专家之前等待路由器输出,导致缓存利用率低下,并且在紧张的显存预算下无法将传输与计算重叠。为了解决这些挑战,我们提出了Mira,一种算法-系统协同设计,使得在单个GPU上能够进行高容量MoE推理。Mira通过将预测性专家管理与定制量化格式相结合,从被动姿态转变为主动姿态。它引入了轻量级的逐层预测器,提前两层预测专家使用情况,从而实现主动预取。这些预测输入到一个由令牌级路由遥测管理的两级HOT+STAGE GPU缓存中,以保留频繁使用的专家,同时分阶段加载预测的专家。为了最小化传输开销,Mira实现了对专家参数的自定义压缩,减少了元数据并提高了打包效率,同时将精度损失降至最低。Mira作为一个完全集成的运行时实现,协调预测器、缓存策略和量化传输,以最大化通信与计算之间的重叠。我们的实验表明,Mira减少了专家引起的停顿。与最先进的基线相比,Mira在内存受限的GPU上实现了平均吞吐量5.71倍的加速。它将首令牌时间加速了11.71倍,并在束搜索推理中实现了3.84倍的平均加速,展示了其在多种推理场景中的有效性。
英文摘要
Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert parameters dominate memory, and token-level routing is dynamic, unpredictable, and skewed. Prior work using offloading and caching remains fundamentally reactive, as systems wait for router outputs before moving experts, leading to inefficient cache utilization and an inability to overlap transfers with compute under tight VRAM budgets. To address these challenges, we propose Mira, an algorithm-system co-design that enables high-capacity MoE inference on a single GPU. Mira shifts from a reactive to a proactive stance by coupling predictive expert management with a tailored quantization format. It introduces lightweight per-layer predictors that anticipate expert usage two layers ahead, enabling proactive prefetching. These predictions feed a two-tier HOT+STAGE GPU cache managed by token-level routing telemetry to retain frequently used experts while staging predicted ones. To minimize transfer overhead, Mira implements a custom compression for expert parameters, which reduces metadata and improves packing efficiency, while minimally degrading accuracy. Mira is implemented as a fully integrated runtime that coordinates predictors, caching policies, and quantized transfers to maximize overlap between communication and compute. Our experiments show that Mira reduces expert-induced stalls. Compared against state-of-the-art baselines, Mira achieves a 5.71x speedup in average throughput on a memory-constrained GPU. It accelerates Time-to-First-Token by 11.71x and achieves a 3.84$x average speedup in beam search inference, demonstrating its effectiveness across diverse inference scenarios.
发表机构
- University of Maryland, College Park(马里兰大学学院公园分校)
机构由 AI 辅助整理,请以论文原文为准。