BASE:基于预测移除误差的批量感知专家选择,用于高效MoE解码
BASE: Batch-Aware Selection of Experts Using Predicted Removal Error for Efficient MoE Decoding
- College of Connected Computing, Vanderbilt University(范德堡大学互联计算学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对MoE批量解码中专家权重传输瓶颈,提出基于预测移除误差的批量感知专家选择方法,训练轻量线性预测器并开发定制内核,无需重训即在多种架构上显著改善质量-效率权衡。
AI中文摘要:
大型语言模型的服务成本日益高昂。在大规模服务系统中,自回归解码的瓶颈往往在于将模型权重从加速器高带宽内存传输到片上SRAM。混合专家(MoE)模型通过仅为每个令牌激活一小部分专家来减少计算量,但这种稀疏性并不能直接转化为批量解码的优势。不同的请求会选择不同的专家;因此,在众多并发请求中,合并后的活跃专家集合可能覆盖专家池的很大一部分,从而需要传输显著更多的专家权重。大多数专家缩减技术独立地为每个令牌做出保留决策,因此无法解决这种批量层面的扩展问题。近期,批量感知方法尝试协调并发请求中的专家使用,并复用已为批次获取的专家。然而,其选择标准主要基于路由器排名或在校准期间收集的专家统计信息。因此,这些标准与因丢弃专家而导致的输出误差并无直接关联,也无法捕捉推理时专家贡献随令牌的变化。我们转而根据移除某个专家会如何改变MoE层输出来对专家进行排序。为了在服务期间应用这一标准,我们在校准期间训练了一个轻量级线性预测器,用于估计每个传入令牌的专家移除成本,并开发了用于成本预测和专家选择的定制GPU内核。在三种MoE架构上,BASE无需重新训练即可改善质量-效率权衡。在Qwen3-30B-A3B上,在严格的专家预算下,与吞吐量相当的最强基线相比,BASE将平均准确率提高了29.5个百分点。在更高的专家预算下,其速度比密集推理快60%,同时准确率差距保持在0.4个百分点以内。
英文摘要:
Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferring model weights from accelerator high-bandwidth memory into on-chip SRAM. Mixture-of-experts (MoE) models reduce computation by activating only a small subset of experts per token, but this sparsity does not translate directly to batched decoding. Different requests select different experts; therefore, the combined active set across many concurrent requests can span a substantial fraction of the expert pool and require significantly more expert weights to be transferred. Most expert-reduction techniques make retention decisions independently for each token and therefore do not address this batch-level expansion. More recently, batch-aware methods have attempted to coordinate expert use across concurrent requests and reuse experts already fetched for the batch. Yet their selection criteria are based primarily on router rankings or expert statistics collected during calibration. Consequently, these criteria are not directly tied to the output error caused by dropping an expert, nor do they capture how an expert's contribution changes across tokens at inference time. We instead rank experts according to how much their removal would change the MoE-layer output. To apply this criterion during serving, we train a lightweight linear predictor during calibration that estimates the expert removal cost for each incoming token, and develop custom GPU kernels for cost prediction and expert selection. Across three MoE architectures, BASE improves the quality-efficiency tradeoff without retraining. On Qwen3-30B-A3B, it improves average accuracy by 29.5 points over the strongest baseline at comparable throughput under a tight expert budget. At a higher expert budget, it is 60% faster than dense inference while remaining within 0.4 accuracy points.