EStream:通过专家虚拟化在移动NPU上实现快速且内存高效的MoE预填充
EStream: Fast and Memory-Efficient MoE Prefill through Expert Virtualization on Mobile NPUs
浏览论文内容
中文总结 AI 辅助
EStream通过专家虚拟化和硬件感知配置,在移动NPU上实现MoE预填充的快速与内存高效,相比基线显著加速并降低内存。
中文摘要 AI 辅助
移动设备厂商和应用程序开发者越来越多地在智能手机上部署大语言模型(LLM),以提供多样化的预填充(prefill)服务。然而,当前系统主要依赖密集模型,其规则计算能高效映射到移动NPU,而能力更强的混合专家(MoE)模型则未被充分利用。MoE预填充不适合移动NPU:NPU图在编译时固定,而MoE在运行时选择专家;且一个请求会触及大多数专家,超过手机内存所能容纳。我们提出EStream,通过分离NPU必须固定的内容与MoE在运行时决定的内容来解决这两个问题。一个编译好的专家图服务于所有专家,每个专家的路由令牌和权重地址在调用时绑定,因此动态MoE执行完全在NPU上进行,无需填充或CPU/GPU回退。专家虚拟化将专家池存储在UFS闪存中,并通过固定大小的NPU可寻址竞技场(arena)逐组进行分页,加载隐藏在计算之后,因此内存受竞技场大小限制而非模型大小。它还引入了一种硬件感知的配置算法,自动配置UFS-NPU流水线并最大化加载与计算的重叠。在涵盖三个7B-16B MoE模型和256-4096令牌提示的18种对比设置中,我们在商用骁龙(Snapdragon)智能手机上评估了EStream。与每种设置下最快的基线相比,EStream实现了2.25-27.57倍的纯预填充TTFT加速,并将峰值物理内存减少了1.19-12.29倍。EStream还能扩展到参数高达467亿的MoE模型。
英文摘要
Mobile vendors and application developers increasingly deploy LLMs on smartphones for diverse prefill-only services. Yet current systems rely mainly on dense models whose regular computation maps efficiently to mobile NPUs, leaving more capable MoEs underused. MoE prefill does not fit mobile NPUs: NPU graphs are fixed at compile time, yet MoE picks experts at runtime; and one request touches most experts, more than a phone can hold in memory. We present EStream, which resolves both by separating what the NPU must fix from what MoE decides at runtime. A single compiled expert graph serves every expert, with each expert's routed tokens and weight address bound at call time, so dynamic MoE execution runs entirely on the NPU without padding or CPU/GPU fallback. Expert virtualization keeps the expert pool in UFS flash storage and pages it through a fixed-size NPU-addressable arena, group by group, with loading hidden behind computation, so memory is bounded by the arena rather than by the model. It further introduces a hardware-aware configuration algorithm that automatically configures the UFS--NPU pipeline and maximizes loading--computation overlap. Across 18 comparative settings covering three 7B--16B MoEs and 256--4,096-token prompts, we evaluate EStream on a commercial Snapdragon smartphone. Compared to the fastest baseline at each setting, EStream achieves a 2.25--27.57X pure-prefill TTFT speedup and reduces peak physical memory by 1.19--12.29X. EStream further scales to MoE models with up to 46.7B parameters.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- Cloud Computing Research Institute, China Telecom(中国电信云计算研究院)
- Temple University(天普大学)
机构由 AI 辅助整理,请以论文原文为准。