发表机构
Holocron Security, Inc(全息安全公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究局部MoE推理在受限内存机器上的可行性,提出MawForge方法,通过磁盘存储模型、按需实例化专家张量等实现,发现其可作为有界执行机制,但非缓存最大化策略,性能受多种因素影响。
AI 中文摘要
稀疏专家混合(MoE)语言模型将总参数数量与每个令牌的活跃计算分开,但本地推理系统通常仍需要完整模型、键值缓存、运行时缓冲区和操作系统空间才能装入快速内存。MawForge测试了一种不同的系统假设:通过将完整模型存储在磁盘上,保留常见张量,并根据需要将路由的专家张量实例化到有界执行缓存中,在受限的统一内存机器上进行局部MoE服务可以变得切实可行。核心发现是,MawForge作为局部MoE推理的有界执行机制和测量基础是有效的,但不是作为缓存最大化策略。性能取决于在专家重用与驻留占用空间、键值缓存大小、量化、路由局部性和macOS内存压力之间进行平衡。
英文摘要
Sparse Mixture-of-Experts (MoE) language models separate total parameter count from per-token active computation, but local inference systems often still require the full model, key-value cache, runtime buffers, and operatingsystem headroom to fit in fast memory. MawForge tests a different systems hypothesis: local MoE serving can be made practical on constrained unified-memory machines by storing the full model on disk, keeping common tensors resident, and materializing routed expert tensors into a bounded execution cache on demand. The central finding is that MawForge is effective as a bounded execution mechanism and measurement substrate for local MoE inference, but not as a cache-maximization policy. Performance depends on balancing expert reuse against resident footprint, KV-cache size, quantization, route locality, and macOS memory pressure.
Comments7 pages