AI 中文总结
SeqMoE通过预测性内存管理和图兼容卸载运行时,在45%专家驻留下实现96.97%命中率与80.22%全负载性能,显著提升MoE卸载效率。
AI 中文摘要
混合专家(MoE)模型为卸载创造了结构性优势:只有一小部分被激活的专家需要驻留在设备内存中,如果它们能够及时加载以供计算,卸载原则上可以接近全负载性能,即所有模型权重都驻留在设备内存中。然而,将MoE的结构性优势转化为实际的卸载收益仍然具有挑战性。我们提出SeqMoE来弥合这一差距。为了最大化专家命中率,我们构建了预测性内存管理:(i)序列到序列预测。我们首次将专家激活预测重新构建为序列建模,实现准确的多步、多层预测,为下游决策提供长而可靠的窗口。(ii)联合预取调度。我们将预取调度表述为具有截止日期的作业排序问题,以最大化预期专家命中率并提高带宽效率。(iii)预测驱动的缓存。利用序列建模的递归特性,我们引入了一种概率性Belady策略用于未来感知的驱逐。为了消除执行瓶颈,我们开发了(iv)图兼容的卸载运行时。我们推导出通用运行时原则,包括计算透明的专家放置和无同步的编排纪律,以实现端到端的图捕获。在45%的专家驻留率下,SeqMoE平均达到96.97%的命中率和80.22%的全负载性能,推进了MoE卸载的最新技术水平。
英文摘要
Mixture-of-Experts (MoE) creates a structural advantage for offloading: only a small fraction of activated experts need to reside in device memory, and if they can be loaded in time for computation, offloading can in principle approach full-load performance, where all model weights reside in device memory. Yet translating MoE's structural advantage into practical offloading gains remains challenging. We propose SeqMoE to bridge this gap. To maximize expert hits, we build predictive memory management: (i) Sequence-to-sequence prediction. We are the first to recast expert activation prediction as sequence modeling, enabling accurate multi-step, multi-layer forecasts that provide a long and reliable window for downstream decisions. (ii) Joint prefetch scheduling. We formulate prefetch scheduling as Job Sequencing with Deadlines to maximize expected expert hits and improve bandwidth efficiency. (iii) Forecast-driven caching. Leveraging the recursive nature of sequence modeling, we introduce a probabilistic Belady policy for future-aware eviction. To eliminate execution bottleneck, we develop (iv) Graph-compatible offloading runtime. We derive general runtime principles encompassing compute-transparent expert placement and synchronization-free orchestration disciplines for end-to-end graph capture. With 45% expert residency, SeqMoE averages a 96.97% hit rate and 80.22% of full-load performance, advancing the state of the art in MoE offloading.