arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OLED-MoE:通过迭代间局部性感知专家卸载加速基于MoE的dLLM推理

OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading

Jingyuan Xiao, Jiayue Wang, Yitao Hu, Xinning Wang, Shi Chen, Ziqi Gong, Zhengchao Wang, Guotao Yang, Sheng Chen, Keqiu Li

arXiv 2609.33385首次发表:更新:

发表机构

Tianjin University(天津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散大语言模型(dLLM)中MoE专家卸载延迟高的问题,提出OLED-MoE,利用迭代间专家重用和置信度预测保留高价值专家,协同CPU-GPU执行,显著降低TPOT并提高缓存利用率。

AI 中文摘要

半自回归扩散大语言模型(dLLM)通过迭代式块状去噪提高解码并行性,但使用混合专家(MoE)层进行扩展会引入大量专家参数,超出内存受限的GPU容量。专家卸载是一种自然的解决方案,然而现有的MoE服务系统针对自回归解码,并依赖迭代内逐层预取:在计算某一层时,预测并加载后续层的专家。在dLLM推理下,块状路由扩大了每次迭代内的活跃专家工作集,使得此类预取难以按时完成,且预测错误时代价高昂。因此,现有的基于预取的解决方案往往退化为按需加载专家,导致解码延迟高。我们提出OLED-MoE,一种专家卸载系统,将优化目标从迭代内预取转向迭代间专家保留。其关键洞察是相邻去噪迭代表现出强烈的专家路由重叠,而令牌置信度指示哪些专家可能被重用。OLED-MoE使用置信度引导的迭代间预测,在GPU内存中保留高价值专家,而不引入额外的预取流量。它进一步通过CPU-GPU协同执行补偿不可避免的缓存未命中,共同考虑动态专家计算负载和预测的未来重用。在多样化的dLLM工作负载中,与最先进的卸载系统相比,OLED-MoE将每个输出令牌的时间(TPOT)降低了1.23倍至7.93倍,并将专家缓存利用率提高了1.44倍至4.23倍。值得注意的是,OLED-MoE在仅使用40%的专家GPU内存的情况下接近全驻留性能,尽管专家内存占用减少了60%,但TPOT仅增加23%。OLED-MoE的源代码在此https URL公开可用。

英文摘要

Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained GPU capacity. Expert offloading is a natural remedy, yet existing MoE serving systems target autoregressive decoding and rely on intra-iteration layer-wise prefetching: while computing one layer, they predict and load experts for subsequent layers. Under dLLM inference, block-wise routing expands the active expert working set within each iteration, making such prefetches difficult to complete in time and costly when mispredicted. Consequently, existing prefetch-based solutions often degenerate into on-demand expert loading with high decoding latency. We propose OLED-MoE, an expert offloading system that shifts the optimization target from intra-iteration prefetching to inter-iteration expert retention. Its key insight is that adjacent denoising iterations exhibit strong expert routing overlap, and token confidence indicates which experts are likely to be reused. OLED-MoE uses confidence-guided inter-iteration prediction to retain high-value experts in GPU memory without introducing extra prefetch traffic. It further compensates unavoidable cache misses through CPU-GPU cooperative execution, jointly considering dynamic expert computation load and predicted future reuse. Across diverse dLLM workloads, OLED-MoE reduces time per output token (TPOT) by 1.23x-7.93x and improves expert cache utilization by 1.44x-4.23x over state-of-the-art offloading systems. Notably, OLED-MoE approaches full-residency performance while using only 40% of the expert GPU memory, incurring merely 23% higher TPOT despite a 60% reduction in expert memory footprint. OLED-MoE's source code is publicly available at https://github.com/flashserve/OLED-MoE.

CommentsAccepted by EuroSys '27 spring. Code available at https://github.com/flashserve/OLED-MoE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑