arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

轨迹驱动评估何时会误导MoE专家缓存?重放语义、工作负载污染与操作模式

Reproducible Evaluation of MoE Expert Caching: Replay Semantics, Workload Contamination, and Operating Regimes

Yu Zhang

arXiv 2608.07911首次发表:更新:

发表机构

China National Chemical Equipment Co. Ltd.(中国化工装备有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对MoE模型专家缓存的轨迹驱动评估,发现重放语义、工作负载污染、操作模式三类因素会误导评估结论,修正后因果下一次使用预测器可部分缩小与离线最优的缓存差距。

AI 中文摘要

混合专家(MoE)模型的规模已超出加速器内存容量,将专家权重卸载至主机内存现已成为标准做法,这使得专家缓存管理成为极具吸引力的优化手段:若某策略能提升缓存命中率,将可降低每个token的专家通信量。但评估该策略的过程是一项测量问题,我们发现此类测量结果存在脆弱性。我们针对三个MoE模型(分别含40、64、128个专家),采用轨迹驱动的事件原子级模拟器,分离出三个会改变结论而非仅改变数值的评估维度:1.重放语义:在融合事件流量约定下,不一致的逐访问重放会使基于近期性的策略的命中率提升27%-29%,而基于频率的策略和静态策略的变化幅度则在4%以内,这会反转策略的排名;2.工作负载污染:使用每类一个指令模板的探测集会生成完全相同的生成前缀,配对渲染干预可使测得的早期窗口效应移动19.4-31.9个百分点,并反转哪些工作负载看起来最适合缓存;3.操作模式:归一化缺失率无法在不同模型间迁移,因此必须报告每步专家集合相对于每一层容量的情况——但仅对相同事件流的时间顺序进行置换,就会使离线最优差距从44.9%变为30.8%,仅靠此方法是不够的。经修正后,与离线最优值的稳定差距仍存在(在13种冻结工作负载组合下为44.2%-45.9%);强制准入预言机将该差距的84.3%-96.6%归因于知晓未来最久才会被使用的驻留专家;作为驱逐规则的因果下一次使用预测器可弥补-11.4%的差距,它在3.4%的情况下能选出最优受害者,而随机驻留块为2.4%,LRU和LFRU则为20.6%-22.1%。我们的观点明确:在我们评估的设置中,较大的离线最优差距会显著高估代表性轻量级因果机制所能弥补的收益。

英文摘要

Mixture-of-Experts (MoE) models have outgrown accelerator memory, and offloading expert weights to host memory is now standard. This makes expert cache management an attractive lever: a policy that raised the hit rate would cut expert traffic per token. Evaluating that is a measurement problem, and we find the measurement fragile. With a trace-driven, event-atomic simulator over three MoE models (40, 64, 128 experts), we isolate three evaluation axes that change conclusions, not just numbers. Replay semantics: under a fused-event traffic contract, an inconsistent per-access replay inflates recency-based policies by 27-29% while leaving frequency-based and static ones within 4%, inverting the policy ranking. Workload contamination: probe sets using one instruction template per category produce verbatim-identical generation prefixes; a matched-pair rendering intervention moves the measured early-window effect by 19.4-31.9 points and reverses which workloads look most cache-friendly. Operating regimes: normalized miss fractions do not transfer across models, so the per-step expert union relative to per-layer capacity must be reported -- yet permuting only the temporal order of an identical event stream moves the offline-optimal gap from 44.9% to 30.8%, so it is not sufficient. Corrected, a stable gap to the offline optimum remains (44.2-45.9% over 13 frozen workload compositions). A forced-admission oracle attributes 84.3-96.6% of it to knowing which resident expert is used furthest in the future. A causal next-use predictor, used as an eviction rule, recovers -11.4% of the gap; it picks an optimal victim 3.4% of the time, against 2.4% for a random resident block and 20.6-22.1% for LRU and LFRU. Our position is narrow: in our evaluated settings a large offline-optimal gap substantially overstates the gains recovered by representative lightweight causal mechanisms.

Comments39 pages, 4 figures, 21 tables. Measurement and scoped negative-results study. Artifacts and SHA-256 manifests: https://doi.org/10.5281/zenodo.21788820 and https://github.com/shijiuzhang/moe-cache-eval. v3: title changed; the generative-AI declaration now names both systems used. v4: metadata only; manuscript unchanged from v3

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑