SAEM:面向思维链推理的内存高效混合专家模型推理的阶段感知型专家管理
SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning
- School of Computing, National University of Singapore(新加坡国立大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对思维链推理中MoE模型的内存与延迟问题,提出SAEM阶段感知型专家管理运行时,通过阶段感知缓存等技术提升吞吐量,在受限GPU内存下平均实现1.33倍的性能提升。
AI中文摘要:
思维链(CoT)提示通过将复杂问题分解为中间步骤提升大语言模型(LLM)的推理能力,但其顺序特性会增加解码延迟与内存占用。混合专家模型(MoE)通过稀疏的专家激活扩展容量,然而其完整的专家权重往往超出GPU内存,需要代价高昂的GPU-CPU数据传输。现有运行时将所有token同等对待,忽略了思维链轨迹的关键结构属性:连续的推理阶段呈现出连贯且可预测的专家激活模式。忽视这种阶段级规律会导致缓存效率低下与不必要的数据移动。本文提出SAEM,一种阶段感知型MoE推理运行时,可检测推理阶段边界并利用阶段级激活连贯性指导专家部署。SAEM结合了阶段感知缓存、专家对齐的token重打包以及原位CPU执行,以减少数据传输与内核碎片化。在数学与科学推理工作负载上,当GPU内存受限时,SAEM相比最强的现有缓存与卸载基线实现了平均1.33倍的吞吐量提升;当校准数据与工作负载匹配时,提升可达1.54倍,证明了面向思维链推理的阶段感知、位置驱动型MoE推理的有效性。
英文摘要:
Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases decoding latency and memory usage. Mixture-of-Experts (MoE) models scale capacity through sparse expert activation, yet their full expert weights often exceed GPU memory and require costly GPU-CPU transfers. Existing runtimes treat all tokens uniformly, overlooking a key structural property of CoT traces: consecutive reasoning stages exhibit coherent and predictable expert activation patterns. Ignoring this stage-level regularity leads to inefficient caching and unnecessary data movement. We propose SAEM, a stage-aware MoE inference runtime that detects reasoning stage boundaries and exploits stage-level activation coherence to guide expert placement. SAEM combines stage-aware caching, expert-aligned token repacking, and in-situ CPU execution to reduce data transfer and kernel fragmentation. On mathematical and scientific reasoning workloads, SAEM achieves an average 1.33x throughput improvement over the strongest state-of-the-art caching and offloading baselines under constrained GPU memory, rising to 1.54x when calibration data matches the workload, demonstrating the effectiveness of stage-aware, locality-driven MoE inference for CoT reasoning.