发表机构
Huazhong University of Science and Technology; AIDX TECH PTE. LTD.; National University of Singapore(华中科技大学; AIDX科技私人有限公司; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对表示工程直接应用于MoE模型的结构不匹配问题,提出与路由无关的RARE框架,通过投影扰动到路由矩阵零空间并修正路由漂移,在三类调控任务中实现了更优的有效性-效用权衡。
AI 中文摘要
表示工程通过修改中间隐藏状态提供了一种轻量级的语言模型行为控制方法,但其直接应用于混合专家(Mixture-of-Experts,MoE)模型会引入结构不匹配。我们首先通过一系列实证研究验证了这种失效模式,发现保留干净的路由能大幅恢复调控性能,且在受控内容下,路由对语义内容的敏感性高于行为变化。受这些发现的启发,我们提出了RARE,一种适用于MoE语言模型的与路由无关的表示工程框架。RARE将任意行为扰动投影到路由矩阵的零空间,从而去除路由可见的组件,并进一步修正传播到选定下游层的路由漂移。为确定该框架中最佳的扰动估计器,我们在三个调控场景(有害性、真实性和事实编辑)下,对六个异构开放权重MoE模型评估了五个估计器。在有害性调控任务中,RARE达到了53.3%的平均攻击成功率,同时保留了67.8%的MMLU准确率,相比基线实现了更强的综合有效性-效用权衡。它还将平均TruthfulQA MC1准确率从41.0%提升至58.6%,将CounterFact效能从16.8%提升至96.3%。这些结果表明,路由一致性是使表示工程适配MoE模型的重要架构考量因素。
英文摘要
Representation engineering offers a lightweight means of controlling language-model behavior by modifying intermediate hidden states, but its direct application to Mixture-of-Experts (MoE) models introduces a structural mismatch. We first verify this failure mode through a series of empirical studies and find that preserving clean routing substantially recovers steering performance and that routing is more sensitive to semantic content than to behavioral changes under controlled content. Motivated by these findings, we introduce RARE, a router-agnostic representation engineering framework for MoE language models. RARE projects arbitrary behavioral perturbations onto the null space of the router matrix, thereby removing router-visible components, and further corrects routing drift propagated to selected downstream layers. To decide the best perturbation estimator in this framework, we evaluate five estimators on six heterogeneous open-weight MoE models across three steering scenarios: harmfulness, truthfulness, and factual editing. On harmfulness steering, RARE reaches an average attack success rate of 53.3% while retaining 67.8% MMLU accuracy, yielding a stronger aggregate effectiveness--utility trade-off than baselines. It further improves average TruthfulQA MC1 accuracy from 41.0% to 58.6% and CounterFact efficacy from 16.8% to 96.3%. These results support routing consistency as an important architectural consideration for adapting representation engineering to MoE models.
Comments20 pages, 3 figures. Paper accepted to the Actionable Interpretability Workshop at COLM 2026