发表机构
AutoArk(AutoArk)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Edge0通过预路由器提前预测路由,实现从SSD流式服务35B MoE模型,在24GB内存下达到20tok/s,并利用恢复LoRA弥补量化损失。
AI 中文摘要
混合专家(MoE)模型在消费级硬件上的推理受限于权重内存:一个35B级别的模型在4位量化下占用19.5GB,而稀疏性减少了每个token的计算量,却并未减少必须持有的字节数。简单地将权重卸载到SSD本身并无帮助,因为第N+1层的专家必须在第N层的输出产生之前就被选定,因此读取操作无法提前开始以隐藏在计算之后。我们提出了Edge0,一个流式MoE推理引擎,通过一个预路由器(prerouter)来弥合这一差距:每层的一个头部提前一个token预测下一层的路由,并且该预测被直接用作路由本身,因此分阶段的专家集合等于路由后的集合,不会丢弃任何内容。一个在student路径上训练、未合并的恢复LoRA,弥补了因int4量化和路由替换而损失的质量。在单台24GB机器上,Edge0在3GiB峰值活跃内存内以20tok/s的速度服务35B MoE模型,在五个公开基准上平均仅比其fp16教师模型低几分。一个8B级别的模型在同一框架上运行,且该框架、检查点和适配器均已开源。
英文摘要
Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's experts must be chosen before layer N's output exists, so the reads cannot start early enough to hide behind compute. We present Edge0, a streaming MoE inference engine that closes the gap with a prerouter: a per-layer head predicts the next layer's routing one token ahead, and the prediction is consumed as the routing itself, so the staged expert set equals the routed set and nothing is dropped. An unmerged recovery LoRA, trained on the student path, pays back the quality lost to int4 quantization and routing replacement. On a single 24GB machine, Edge0 serves a 35B MoE at 20tok/s inside 3GiB of peak active memory, within a few points of its fp16 teacher on average across five public benchmarks. An 8B tier runs on the same framework, and the framework, checkpoints, and adapters are open source.