发表机构
Seoul National University; Neural Processing Research Center(首尔大学; 神经处理研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Q-Strata是一种双层分配器,通过为混合专家MoE大型语言模型的每个模块设置一个预算,优化模型级目标以实现混合精度量化,在低比特 regime 下优于GPTQ、MxMoE和GEMQ。
AI 中文摘要
混合精度量化(MPQ)为大型语言模型(LLM)的每个线性层分配不同的比特宽度,以在固定预算下最小化量化导致的质量损失,但混合专家(MoE)模型的每个MoE模块的每个专家中都包含这些线性层,因此其分配空间远大于密集模型。现有方法要么在每个模块内以统一的模块预算进行分配,要么通过加性代理在模块间进行分配,均未直接优化耦合模块选择的模型级目标。我们提出Q-Strata,一种双层分配器,其用廉价代理对模块内分配进行排序,并通过组装后的量化模型评估的模型级目标在模块间进行分配。其内部阶段在精细间隔的预算下缓存每个模块的候选帕累托前沿,使外部阶段为每个模块设置一个预算,而非为每个线性层设置比特宽度。搜索简化为每个模块一个预算后,外部阶段直接优化该模型级目标,捕获加性代理缺失的模块间耦合。在Mixtral-8x7B-Instruct、Qwen1.5-MoE-A2.7B和DeepSeek-V2-Lite上,Q-Strata在低比特 regime 下始终比均匀比特宽度的GPTQ及最先进的MoE MPQ方法MxMoE、GEMQ实现更低的WikiText2困惑度。代码可在该https URL获取。
英文摘要
Mixed-precision quantization (MPQ) assigns a different bitwidth to each linear layer of a large language model (LLM) to minimize the quantization-induced quality loss under a fixed budget, but Mixture-of-Experts (MoE) models contain these layers in every expert of every MoE block, so the allocation space grows far larger than in a dense model. Existing methods either allocate within each block under a uniform per-block budget, or allocate across blocks through an additive proxy, and neither directly optimizes a model-level objective over the choices that couple the blocks. We propose Q-Strata, a bi-level allocator that ranks within-block assignments with a cheap proxy and allocates across blocks with a model-level objective evaluated on the assembled quantized model. Its inner stage caches a Pareto frontier of candidates per block over finely spaced budgets, leaving the outer stage to set one budget per block instead of a bitwidth for every linear layer. With the search reduced to one budget per block, the outer stage optimizes this model-level objective directly, capturing the inter-block coupling that additive proxies miss. On Mixtral-8x7B-Instruct, Qwen1.5-MoE-A2.7B, and DeepSeek-V2-Lite, Q-Strata consistently achieves lower WikiText2 perplexity than uniform-bitwidth GPTQ and the state-of-the-art MoE MPQ methods MxMoE and GEMQ in the low-bit regime. The code is available at https://github.com/snu-mllab/Q-Strata/tree/main.
CommentsEMNLP 2026 Long Paper - Main Conference