发表机构
Hainan University; Xiamen University; Zhejiang Normal University(海南大学; 厦门大学; 浙江师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ITC-MoE提出重要性引导的令牌感知压缩框架,通过自适应Tucker分解和令牌级补偿路由,在30%压缩预算下保持96.33%准确率并实现7.22倍加速,降低MoE扩散语言模型的计算存储成本。
AI 中文摘要
混合专家(MoE)扩散语言模型(DLMs)提供了灵活的并行解码和更大的模型容量,但其大量的专家参数带来了可观的计算和存储成本。现有的低秩MoE压缩方法主要依赖于静态分解和固定秩分配,忽略了MoE DLMs的独特性质。具体而言,我们识别出两个性质:跨模态非均匀冗余,即参数冗余和对秩截断的敏感性在输入、输出和专家模态之间有所不同;以及令牌级利用差异,即热令牌和冷令牌表现出不同的频谱特征和专家激活模式。为了应对这些挑战,我们提出了ITC-MoE,一个面向MoE DLMs的重要性引导的令牌感知压缩框架。ITC-MoE由两个互补组件组成。首先,重要性引导的自适应Tucker压缩(IATC)将激活和梯度重要性纳入专家权重变换,跨多个模态联合分解专家权重,并在固定参数预算下自适应分配秩。其次,令牌感知补偿与路由(TCR)对压缩敏感的热令牌应用轻量级低秩补偿,并限制具有集中路由模式的冷令牌的候选专家集。通过将压缩容量和推理执行同时适应参数冗余和令牌级变化,ITC-MoE大幅降低了MoE DLMs的计算和存储成本,同时保持了其生成质量。例如,在SDAR-30B-A3B-Chat-b32上,ITC-MoE在30%压缩预算下在MultiArith上保持了96.33%的准确率,同时实现了高达7.22倍的端到端加速。代码可在以下网址公开获取:此HTTPS URL。
英文摘要
Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fixed rank allocation, which overlook the distinctive properties of MoE DLMs. Specifically, we identify two properties: cross-mode non-uniform redundancy, where parameter redundancy and sensitivity to rank truncation vary across the input, output, and expert modes, and token-wise utilization variation, where hot and cold tokens exhibit distinct spectral characteristics and expert activation patterns. To address these challenges, we propose ITC-MoE, an Importance-guided Token-aware Compression framework for MoE DLMs. ITC-MoE consists of two complementary components. First, Importance-guided Adaptive Tucker Compression (IATC) incorporates activation and gradient importance into expert weight transformation, jointly factorizes expert weights across multiple modes, and adaptively allocates ranks under a fixed parameter budget. Second, Token-aware Compensation and Routing (TCR) applies lightweight low-rank compensation to compression-sensitive hot tokens and restricts the candidate expert set for cold tokens with concentrated routing patterns. By jointly adapting compression capacity and inference execution to both parameter redundancy and token-wise variation, ITC-MoE substantially reduces the computation and storage costs of MoE DLMs while preserving their generation quality. For example, on SDAR-30B-A3B-Chat-b32, ITC-MoE maintains an accuracy of 96.33% on MultiArith under a 30% compression budget, while achieving up to a 7.22x end-to-end speedup. The code is publicly available at https://github.com/lianjunl13-sudo/ITC-MoE.