AI 中文总结
本文通过在XLCoST基准上对Qwen3.6-35B-A3B模型的逐层敏感性分析,发现MoE层敏感性具深度依赖性,提出聚焦晚期层的掩码策略,可在保留输出质量的同时减少专家数量,为MoE模型压缩提供了实证基础与实用路径。
AI 中文摘要
混合专家(MoE)架构通过稀疏激活在保留计算效率的同时扩展了大语言模型(LLMs)的规模。尽管MoE架构已被广泛采用,但各MoE层的相对重要性仍未得到充分表征,尤其在模型压缩领域。本文在XLCoST跨语言代码翻译基准上,采用基于幅值的专家掩码,对Qwen3.6-35B-A3B模型(含40个MoE层,每层256个专家,采用top-8路由)开展了系统的逐层敏感性分析。我们在三台H100 GPU服务器上,针对100、300和500个提示评估规模,进行了多阶段研究。核心发现是层敏感性具有强烈的深度依赖性:早期层(0-9)和中间层(10-29)对专家掩码高度脆弱,而晚期层(30-39),尤其是极晚期层(35-39),可容忍对低幅值专家的激进掩码。在300个提示规模下,对所有层应用30%的掩码仅保留150/300个良好+相似输出,而聚焦晚期的掩码策略在掩码640-1145个专家的同时,可保留249-255/300个输出。在后续500个提示的保留验证切片上,针对极晚期层(35-39)应用50%掩码的窄范围策略,在所有测试方案中实现了最强的质量/掩码专家权衡,在掩码10240个总专家中的640个的同时,保留了419/500个良好+相似输出。我们还研究了每个令牌的活跃专家数从8减少到6的top-k路由宽度,在100个提示探针上观察到大幅的 wall-clock 减少,且无良好+相似输出损失,但该减少尚未能与激进专家掩码良好组合。这些发现为深度感知MoE专家掩码提供了实证基础,并为物理权重手术、基于激活的专家评分及基于训练的恢复建立了实用路径。
英文摘要
Mixture-of-Experts (MoE) architectures scale large language models (LLMs) while preserving computational efficiency through sparse activation. Despite their widespread adoption, the relative importance of individual MoE layers remains insufficiently characterized, particularly for model compression. This paper presents a systematic layer-wise sensitivity analysis of the Qwen3.6-35B-A3B model (40 MoE layers, 256 experts per layer, top-8 routing) using magnitude-based expert masking on the XLCoST cross-lingual code translation benchmark. We conduct a multi-phase study spanning 100, 300, and 500 prompt evaluation scales across three H100 GPU servers. Our central finding is that layer sensitivity is strongly depth-dependent: early layers (0-9) and middle layers (10-29) are highly fragile to expert masking, while late layers (30-39), and especially very-late layers (35-39), tolerate aggressive masking of low-magnitude experts. Flat all-layer masking at 30% retains only 150/300 Good+Similar outputs at 300-prompt scale, whereas late-focused policies retain 249-255/300 while masking 640-1,145 experts. On a later 500-prompt held-out validation slice, the narrow very-late policy (layers 35-39 @ 50%) achieves the strongest quality/masked-expert tradeoff among tested candidates, retaining 419/500 Good+Similar outputs while masking only 640 of 10,240 total experts. We additionally characterize top-k routing width reduction from 8 to 6 active experts per token, which shows a large observed wall-clock reduction on a 100-prompt probe with no Good+Similar loss, though it does not yet compose cleanly with aggressive expert masking. These findings provide an empirical foundation for depth-aware MoE expert masking and establish a practical path toward physical weight surgery, activation-based expert scoring, and training-based recovery.