发表机构
Ajou University(亚洲大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Colla-Q,一种基于激活熵的位宽分配框架,通过平衡MoE中各专家的性能来提升量化模型的整体性能并减少校准依赖。
AI 中文摘要
本文提出了一种基于激活熵的混合专家(MoE)量化方法。尽管量化降低了内存和计算成本,但它可能显著降低性能。特别是在量化的MoE模型中,性能下降尤为明显,因为单个专家拥有的参数数量较少,且对低位表示敏感。考虑到MoE作为集成模型运作,由路由专家协作贡献,某个特定专家因量化导致的显著性能下降会损害模型整体性能。因此,我们提出了Colla-Q,一种位分配框架,通过基于激活熵的位宽分配算法来维持各专家间的性能平衡。该方法鼓励每个专家在量化模型中协作运作,从而1)提升MoE整体性能,2)减少对校准数据集的依赖。由于统一调整每个专家的性能有助于提高MoE模型的鲁棒性和稳定性,所提出的MoE量化方法能在不同校准数据集上更一致地泛化。我们的代码可在以下网址获取:此https URL
英文摘要
In this paper, we present a Mixture-of-Experts (MoE) quantization method based on activation entropy. Although quantization reduces memory and computational costs, it can substantially degrade performance. In particular, performance decline is pronounced in quantized MoE models, where individual experts have a small number of parameters that are sensitive to low-bit representation. Considering that MoE operates as an ensemble model with collaborative contributions from routed experts, a significant performance decline of a particular expert due to quantization can harm model performance. Therefore, we propose Colla-Q, a bit-allocation framework to maintain balanced performance across experts through an activation-entropy-based bit-width allocation algorithm. This approach encourages each expert to operate collaboratively in the quantized model, thereby 1) improving the overall MoE performance and 2) reducing the dependence on the calibration dataset. Since uniformly adjusting each expert's performance facilitates robustness and stability of the MoE model, the proposed MoE quantization method can generalize more consistently across different calibration datasets. Our code is available at: https://github.com/mmai-laboratory/Colla_Q
CommentsAccepted by the Conference on Empirical Methods in Natural Language Processing (EMNLP) 2026