arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向内存高效MoE推理的缓存感知联合路由器适配

Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference

Zhenhe Wu, Yaping Jin, Qinghua Xing, Hang Zhou, Wei He, Xianjie Wu, Xianfu Cheng, Jian Yang, Hanting Chen

arXiv 2609.04895首次发表:更新:

发表机构

Huawei Technologies; Beihang University; Tianjin University; University of Sydney; Beijing Information Science & Technology University(华为技术有限公司; 北京航空航天大学; 天津大学; 悉尼大学; 北京信息科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对MoE推理中专家集合超GPU内存导致的重复权重传输问题,提出缓存感知后训练框架,联合适配MoE主干与缓存路由器,在Qwen3、GPT-OSS等模型上提升缓存命中率并减少流量。

AI 中文摘要

混合专家(MoE)模型对每个token仅激活一小部分专家,但完整专家集合常超出GPU内存,导致解码过程中出现重复的权重传输。我们将专家缓存管理建模为模型侧的算法问题,提出一种缓存感知的后训练框架,该框架在推理时保留原生Top-K专家选择规则的同时,联合适配MoE主干网络与轻量级辅助缓存路由器。其仅更新模式「时间路由器」可预测同层复用情况,为后续token保留专家而无需主动加载;完整的「时空路由器」新增空间路由器,利用因果前驱的隐藏状态在目标层访问前优化时间缓存。我们在Qwen3和GPT-OSS模型上,针对GSM8K、MATH和CommonsenseQA数据集评估两种模式:时间路由器相比匹配的仅语言模型基线,持续提升缓存命中率并减少专家权重流量;在Qwen3上,时空路由器在三项任务中实现最佳的负载调整效率,相比评估中最强的预取基线,调整后命中率提升1.15至18.03个百分点,流量降低4.6%至53.3%,在GPT-OSS上的结果具有竞争力但依赖任务类型;仅辅助模块的 ablation 实验保留了基线准确率但仅带来有限的缓存增益,而联合后训练可产生更大改进;敏感性分析显示,缓存容量控制传输需求,优化预算决定访问前覆盖范围与主动流量间的权衡。

英文摘要

Mixture-of-Experts (MoE) models activate few experts per token, yet their full expert sets can exceed GPU memory and require repeated weight transfers during decoding. We formulate expert-cache management as a model-side algorithmic problem and propose cache-aware post-training that jointly adapts the MoE backbone and lightweight auxiliary routers while preserving the native inference-time Top-K rule. The update-only Temporal Router learns same-layer retention across tokens without proactive loading. The full Spatio-Temporal Router adds a Spatio Router that uses the causal predecessor's hidden state to refine the temporal cache before target-layer access. We evaluate both modes on Qwen3 and GPT-OSS across GSM8K, MATH, and CommonsenseQA. Temporal Router consistently improves hit rate and reduces expert-weight traffic over matched LM-only baselines. On Qwen3, the full mode improves adjusted hit rate by 1.15--18.03 points and reduces traffic by 4.6--53.3\% relative to the strongest evaluated prefetching baseline; GPT-OSS results are competitive but task-dependent. Auxiliary-only training preserves baseline accuracy but yields modest coverage gains; joint post-training achieves substantially higher coverage. Sensitivity analyses distinguish the effects of cache capacity, refinement budget, and cache-loss weight on coverage, traffic, and quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑