arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12103cs.OS

谁应掌控专家缓存?万亿参数MoE推理的内核管理分层

Who Should Own the Expert Cache? Kernel-Managed Tiering for Trillion-Parameter MoE Inference

Yuan Si, Yufeng Lin, Daming Li, Jialu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对万亿参数MoE推理的专家缓存,对比用户空间实现与OS页缓存方案,发现内核管理缓存可提升解码性能1.09-1.10倍,提出让内核掌控驱逐、模型知识用于准入与建议的设计原则。

中文摘要 AI 辅助

专家池远超DRAM容量的混合专家(MoE)模型,迫使每个服务系统都配备缓存,但现有系统通常在用户空间实现该缓存,采用以专家为粒度、按频率排序且明确固定的分层。本文研究操作系统(OS)已提供的替代方案:将页缓存作为专家分层。我们使用三个MoE模型的路由跟踪数据,每个层的专家数量从128到896不等,其中包括一个拥有1.45TB专家池的万亿参数生产模型,在GH200节点上针对完整专家池原生重放这些跟踪数据,并通过三种独立机制强制执行容量限制,得出四个主要结果:第一,迭代时间和设备流量是缓存容量的平滑、可复现函数,使DRAM成为万亿参数服务的实用容量调节旋钮;第二,深压力拐点是回收产物,需要MGLRU和气球式大部分锁定内存共同作用,而cgroup限制和物理内存配置不会产生此类放大效应,表明基于气球的研究可能会将压力下的设备流量高估约2倍;第三,在强制等内存墙条件下,未调优的内核LRU的请求命中率与同域的专家频率表(oracle)基本相当,在256GB时分别为75.3%和74.6%,而oracle的机制优势仅为1.09倍,且在异域消失,此时LRU仍保持70%-71%的命中率;第四,实测召回率为64.7%的路由预取作为内核预读建议仅提供0.3%的性能提升,作为同步预取则无提升。整体而言,启用内核管理缓存可将解码性能提升1.09-1.10倍,且在9组平衡对中输出完全相同。由此得出的设计原则很简单:在该场景下,让内核掌控驱逐操作,而模型特定知识最好用于准入控制和预取建议。

英文摘要

Mixture-of-experts models whose expert pools exceed DRAM capacity require a weight-residency tier. Existing systems manage it in user space with expert-granular placement, frequency-based admission, and explicit pinning. We evaluate whether the operating system page cache can instead serve as the expert tier, using router traces from three MoE models with 128 to 896 experts per layer; the trillion-parameter production model's traces are replayed natively against its full 1.45 TB expert pool on GH200 hardware. Capacity is enforced by three independent mechanisms. Iteration time varies smoothly with cache size (run-to-run spread <=4%), and device traffic follows the same trend. Under severe pressure the outcome depends on reclaim: device traffic rises above miss demand only when MGLRU, the tested kernels' default, is combined with balloon-style, mostly mlocked memory, a result reproduced on two machines; cgroup limits and mem= boots show no such behavior, so balloon-based studies can overstate low-capacity device traffic by about 2x. At equal enforced memory, kernel recency serves essentially the same demand as an oracle static-frequency policy computed from the replay trace. In the pread-based replay the oracle-pinned arena stays 1.09-1.11x faster, a gap that is the cost of the page-cache hit and reclaim path, but its static table degrades under domain shift while recency remains stable. At 64.7% measured recall, router lookahead changes median time by 0.3% when delivered as kernel readahead advice; perfect one-layer advice gains 5.0% through the same interface and nothing through blocking reads. End-to-end at ample capacity, enabling page-cache admission speeds steady decode by 1.09-1.10x in a production CUDA engine with token-identical outputs. These measurements favor kernel-managed eviction, with model knowledge applied to admission and predictive advice.

↑