arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

共享低秩基分解用于无数据混合专家模型压缩

Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression

Tianxiao Cao, Jiahe Shao, Yuning Qiu, Kyohei Atarashi, Hisashi Kashima, Qibin Zhao

arXiv 2610.09342首次发表:更新:

发表机构

Kyoto University; The University of Tokyo; RIKEN AIP(京都大学; 东京大学; 理化学研究所革新智能统合研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对MoE大模型压缩,提出共享低秩基分解(SLBF)无数据权重重构方法,通过专家间共享秩-k基降低重构误差,在16B至122B五个架构上优于剪枝、合并和重构三类方法。

AI 中文摘要

混合专家(MoE)大语言模型通过稀疏路由将容量与计算解耦,但其庞大的参数数量带来了存储和服务方面的挑战。我们分析了三类MoE压缩方法:专家剪枝、专家合并和权重重构,并推导了结构误差界,表明剪枝和合并可能产生与路由和专家异质性相关的非消失误差。相比之下,权重重构通过保留专家结构和路由来避免这些结构成本。受此分析启发,我们提出了共享低秩基分解(SLBF),一种无数据的权重重构方法,使用专家间共享的秩-k基,从而实现更丰富的跨专家共享、更快的收敛速度和更低的重构误差。事后规范固定(gauge fixing)在不增加表示成本的情况下移除冗余参数。在跨越160亿到1220亿参数的五个MoE架构上,SLBF consistently优于来自所有三个压缩家族的方法。

英文摘要

Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑