arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15687cs.AI

THESIS-MoE:可训练的分层提取与混合专家模型中谄媚行为的引导

THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts

Kareem Hassani, Chaymaa Abbas, Lama Mawlawi, Mariette Awad

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出THESIS-MoE,通过共享对比信号定位MoE模型中谄媚行为的计算子电路,采用有条件干预方法,在保留知识的同时消除了高达90%的信念诱导谄媚行为。

中文摘要 AI 辅助

谄媚行为是语言模型为匹配用户表述的信念而改变回答的倾向,是一种常见的对齐失败问题。现有激活引导方法通常在模型中均匀应用单一对比方向,属于无条件干预,即使不存在谄媚行为也会改变激活,以牺牲知识保留为代价实现行为修正。在混合专家(Mixture-of-Experts, MoE)模型中,已有研究表明行为仅编码在专家计算而非路由决策中,这使得精准的行为引导极具挑战性。本研究引入一种共享对比信号,该信号由包含和不包含用户表述信念的匹配提示构建,可识别MoE分层结构中谄媚行为的位置,并仅在存在该行为的位置实施干预。我们将定位问题形式化为对MoE块、专家、注意力块及注意力头组成的粒度阶梯的因果搜索,比较了无条件减法与两种有条件替代方案:基于解析投影的减法,以及一种可学习的逐词门控,该门控在保持模型权重冻结的同时引导模型远离谄媚行为。我们在三个MoE模型上进行评估,同时测量谄媚行为、通用知识及推理基准。我们的有条件干预消除了高达90%由信念诱导的谄媚行为。结果表明,谄媚行为存在于可识别的计算子电路中,可被选择性引导,同时保持良好的消除-保留权衡。

英文摘要

Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure. Existing activation steering methods typically apply a single contrastive direction uniformly throughout the model, which is an unconditional intervention that alters activations even when no sycophantic behavior is present, trading knowledge retention for behavioral correction. In Mixture-of-Experts (MoE) models, prior work further suggests that behavior is encoded within expert computations rather than routing decisions alone, making precise behavioral steering particularly challenging. In this work, we introduce a shared contrastive signal, built from matched prompts with and without a stated belief, that identifies where sycophancy lives across the MoE hierarchy and drives interventions that act only where the behavior is present. We formulate localization as a causal search over a granularity ladder of MoE blocks, experts, attention blocks, and heads, and compare unconditional subtraction against two conditional alternatives: an analytic projection-based subtraction and a learned per-token gate that steers the model away from sycophancy while keeping its weights frozen. We evaluate on three MoE models measuring sycophancy alongside general knowledge and reasoning benchmarks. Our conditional interventions removed up to 90\% of the belief-induced sycophancy. Our results demonstrate that sycophancy resides in identifiable computational subcircuits and can be selectively steered while maintaining a favorable removal-retention trade-off.

↑