arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07814cs.LGcs.AI

形状变异专家压缩:LorExperts与BTExperts

Shape Mutating Expert Compression:LorExperts and BTExperts

Inesh Chakrabarti, Sourjya Roy, Bowen Bao, Thiago Crepaldi, Spandan Tiwari, Ashish Sirasao

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对混合专家(MoE)语言模型的专家权重压缩问题,提出LorExperts与BTExperts方法,在保留所有专家与原始路由的前提下实现约50%专家压缩,性能优于基线且随专家数量增长优势更显著。

中文摘要 AI 辅助

混合专家(MoE)语言模型以每token低计算量实现高容量,但低成本部署需压缩其大量专家权重矩阵。专家剪枝(如REAP)与合并可降低成本,但会牺牲精度且需重新训练路由;专家的低秩增量分解(如D^2-MoE)保留所有专家与路由,但随专家数量增长性能急剧下降,因单一共享分量无法近似众多近正交专家。由于MoE专家权重近正交,单一共享分量(如先前增量分解)随专家数量扩展性能差;本文发现专家仍会组织为功能共激活社区,且与权重相似性解耦。基于此,本文提出LorExperts,一种保留路由的压缩方法:将专家聚类,每个聚类保留一个全精度主导专家,其余成员表示为对局部主导专家的低秩修正。LorExperts保留所有专家与原始路由(无需重新训练路由)。在Qwen3-30B-A3B和Gemma-4-26B-A4B上实现约50%专家压缩时,LorExperts在多数任务上的下游精度与困惑度优于基线;与D^2-MoE的差距随专家数量E增大而扩大。本文还给出LorExperts的重建微调过程,以及BTExperts——一种主导专家与修正项的树状组织,支持推理时共享计算的摊销。

英文摘要

Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices. Expert pruning (e.g., REAP) and merging reduce cost but sacrifice accuracy and require retraining the router; low-rank delta decomposition of experts (e.g., D^2-MoE) preserves all experts and the router, but degrades sharply as the expert count grows because a single shared component cannot approximate many near-orthogonal experts. Because MoE expert weights are near-orthogonal, a single shared component (as in prior delta decomposition) scales poorly with the expert count; we show that experts nonetheless organize into functional co-activation communities that are decoupled from weight similarity. Building on this, we introduce LorExperts, a router-preserving compression method that clusters experts, keeps one full-precision dominant per cluster, and represents the remaining members as low-rank corrections to their local dominant. LorExperts retains all experts and the original router (no router retraining). At ~50% expert compression on Qwen3-30B-A3B and Gemma-4-26B-A4B, LorExperts preserves downstream accuracy and perplexity better than the baselines on most of the tasks; the margin over D^2-MoE grows with expert count E. We further give a reconstruction fine-tuning procedure for LorExperts, and BTExperts, a tree organization of dominants and corrections that enables inference-time amortization of shared computation.

发表机构

  • Advanced Micro Devices(超威半导体公司)

机构由 AI 辅助整理,请以论文原文为准。

↑