发表机构
Distiller Labs(Distiller实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现MoE模型对道德内容的编码看似稳健实则脆弱,归因于输出稀释,其稳健性比同等规模密集模型低4.2倍,且该脆弱性为架构性而非学习所得。
AI 中文摘要
混合专家(MoE)模型对道德内容的编码看似和密集模型一样稳健,实则脆弱得多。在OLMoE-1B-7B模型中,线性探针可从几乎每一个专家层组合中恢复道德效价,平均峰值层准确率超过90%。但这些表征在激活噪声水平下会崩溃,而同等规模的密集模型能轻松耐受,稳健性差异达4.2倍。我们将此归因于输出稀释:由于MoE模块在贡献到残差流前会对激活的专家取平均,到达下游层的前馈信号比密集MLP小近两个数量级。我们关注的道德信息在聚合时完整保留,但规模极易被扰动淹没。路由本身在噪声下保持稳定,脆弱性完全源于稀释的聚合结果。检查点轨迹证实这是架构性的,非学习所得:专家从未专业化,准确率在最初数千步内就达到饱和。在稀疏架构中,冗余编码并不等同于稳健编码。
英文摘要
Mixture-of-Experts (MoE) models appear to encode moral content as robustly as dense models, yet prove far more fragile in their encoding. In OLMoE-1B-7B, linear probes recover moral valence from nearly every expert-layer combination, with mean peak-layer accuracy above 90%. But these representations collapse under levels of activation noise that a dense model of matched size easily tolerates, with a 4.2-fold difference in robustness. We trace this to output dilution. Because the MoE block averages across active experts before contributing to the residual stream, the feedforward signal reaching downstream layers is nearly two orders of magnitude smaller than in a dense MLP. Moral information, our interest, survives aggregation intact but at a scale trivially overwhelmed by perturbation. Routing itself remains stable under noise while the vulnerability originates entirely in the diluted aggregate. Checkpoint trajectories confirm this is architectural, not learned. Experts never specialize and accuracy saturates within the first few thousand steps. In sparse architectures, redundant encoding does not imply robust encoding.
Comments18 pages, 4 figures. Code and outputs at https://github.com/deepsteer/deepsteer