arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12550cs.LG

固定量化混合专家实例池上的质量约束路由

Quality-Constrained Routing over a Fixed Pool of Quantized Mixture-of-Experts Instances

  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhenghong Huang, Hongfan Wu, Jiheng Zhang

AI总结:

针对固定量化MoE实例池,提出FWP脆弱性加权困惑度预测请求级质量风险,并通过窗口级线性规划路由,在质量预算下最大化吞吐,较静态W4提升28.4%吞吐。

AI中文摘要:

量化混合专家(MoE)服务可以持有同一基础模型的多个预物化实例,但量化损伤在不同请求和位宽之间差异显著。由于实例物化和副本数量消耗内存且需要缓慢的重新配置,我们将其视为上游供应决策,并在固定驻留池内研究路由问题。在此固定池边界内,我们对每个请求进行路由,以在类别级预期质量退化预算和实测实例容量下最大化建模吞吐量。为预测这种请求特定风险,我们引入FWP(脆弱性加权困惑度),该指标基于参考实例预填充中的提示词元计算,并针对候选实例退化进行校准。FWP的基础是精确的双专家亲和力-脆弱性分解,以及条件多层top-$k$展开,其偏差、交互、路径变化、可分离性和高阶项保持显式。利用这些校准风险,窗口级线性规划产生一个有符号的降奖励分数,该分数在最优价格和原始可行平局分配下与LP最优解满足KKT一致性。在88个扩展Qwen提示上,量化全部6,144个专家块的完整W2、W3和W4实例的平均$\Delta$NLL分别为0.9437、0.1832和0.0513。在同一总体和$\tau=0.1513$下,FWP分配达到1.284倍的离线基于模型的乘数,而请求无关混合为1.253倍,静态W4为1.000倍,相对FWP增益为2.5%。

英文摘要:

Quantized Mixture-of-Experts (MoE) services can hold several pre-materialized instances of one base model, but quantization damage varies sharply across requests and bitwidths. Because instance materialization and replica counts consume memory and require slow reconfiguration, we treat them as upstream provisioning decisions and study routing within a fixed resident pool. Within this fixed-pool boundary, we route each request to maximize modeled throughput under a class-level expected quality-degradation budget and measured instance capacities. To predict this request-specific risk, we introduce FWP (Fragility-Weighted Perplexity), computed from prompt tokens on a reference-instance prefill and calibrated to candidate-instance degradation. Underlying FWP is an exact two-expert affinity--fragility decomposition and a conditional multi-layer top-$k$ expansion whose bias, interaction, route-change, separability, and higher-order terms remain explicit. Using these calibrated risks, a window-level linear program yields a signed reduced-reward score that is KKT-consistent with the LP optimum under optimal prices and primal-feasible tie allocation. On 88 extended Qwen prompts, complete W2, W3, and W4 instances quantizing all 6,144 expert blocks incur mean $Δ$NLL of $0.9437$, $0.1832$, and $0.0513$. Under the same population and $τ=0.1513$, FWP allocation reaches a $1.284\times$ offline model-based multiplier versus $1.253\times$ for request-agnostic mixing and $1.000\times$ for static W4, an incremental $2.5\%$ relative FWP gain.

补充信息

↑