arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18261cs.AIcs.LG

可设计为可缓存的?针对边缘内存带宽瓶颈训练混合专家模型路由器:一项带系统测量研究的预注册负面结果

Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study

  • University of Cumberlands(坎伯兰大学)

机构由 AI 辅助整理,请以论文原文为准。

Shriniwas Ramesh Suram

AI总结:

该研究针对边缘部署MoE模型的内存带宽瓶颈,通过系统测量与预注册实验发现,训练路由器提升缓存性的方案无法满足预注册的困惑度阈值,仅结合无训练重路由可实现高效缺失减少。

AI中文摘要:

在单块8GB GPU上部署参数规模达2350亿的混合专家(MoE)模型时,瓶颈并非计算能力,而是内存带宽:解码阶段必须从存储各专家的层级中流式读取每个token对应的活跃专家,而在消费级硬件中,多数专家存储在比RAM慢得多的SSD上。我们在Qwen3-235B(Q4_K_M,134GB)上量化了该带宽瓶颈:实测解码速度为0.44 token/s(预热后),与每token字节数/带宽模型匹配;而本应分摊一次磁盘扫描开销的批处理方案,在批次大小为32时因分页抖动而崩溃。我们构建了llama-moe-trace这一无侵入式路由器遥测工具,并在Qwen3-30B上测量路由情况:相邻token的专家复用概率为2.0倍,95%的流量使用52.5%的专家,占专家总数13.4%的LRU缓存可满足66%的请求。随后我们探究可缓存性是否可训练:我们预注册了对1.37亿参数MoE语言模型的训练,引入辅助局部性和领域路由器损失,同时设置缓存缺失减少量和困惑度的联合标准。该机制有效(缺失量最多降低60%;静态固定命中率达99%),但所有配置均未通过预注册的困惑度≤1%的阈值——缺失减少量与模型质量紧密耦合。同期StickyMoE研究称,在单领域、参数规模小于2500万的模型上,该损失近乎无开销;但在多领域、1.37亿参数的模型上,我们发现该开销确实存在。我们的贡献包括:这项预注册、更严格标准的多领域评估,以及边缘部署测量。对3.4亿参数模型的验证显示,该开销未随规模缩小(反而略有上升)。我们进一步证明,结合训练的局部性与无训练的缓存感知重路由,在两种规模下均可实现约80%的缺失减少量,且困惑度增幅≤3.4%,比单独使用任一方案成本低得多;而基于领域的预取并无帮助。所有代码、追踪数据和预注册内容均已公开。

英文摘要:

Serving a 235B-parameter Mixture-of-Experts (MoE) model on a single 8 GB GPU is bottlenecked not by compute but by memory bandwidth: decode must stream each token's active experts from whichever tier holds them, and on consumer hardware most experts sit on an SSD far slower than RAM. We quantify this bandwidth wall on Qwen3-235B (Q4_K_M, 134 GB): measured decode is 0.44 tok/s warm, matching a bytes-per-token / bandwidth model, while a batching scheme that should amortize one disk sweep instead collapses at batch 32 from paging thrash. We build llama-moe-trace, a zero-surgery router-telemetry tool, and measure routing on Qwen3-30B: adjacent-token expert reuse is 2.0x chance, 95% of traffic uses 52.5% of experts, and an LRU cache of 13.4% of experts serves 66% of requests. We then ask whether cacheability is trainable: we pre-register training of 137M MoE language models with auxiliary locality and domain router losses, under joint criteria on cache-miss reduction and perplexity. The mechanism works (misses down up to 60%; a 99% static-pin hit rate) but every configuration fails the pre-registered <=1% perplexity gate -- miss reduction and quality are tightly coupled. Concurrent StickyMoE reports the same loss as near-free on single-domain sub-25M models; on multi-domain 137M we find the tax real. Our contribution is this pre-registered, stricter-criterion, multi-domain evaluation plus edge-serving measurements. A 340M rung shows the tax does not shrink with scale (it rises slightly). We further show training-free cache-aware rerouting stacks with trained locality -- together ~80% miss reduction at <=3.4% perplexity at both sizes, far cheaper than either alone -- while domain-primed prefetching does not help. All code, traces, and the pre-registration are released.

补充信息

↑