arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越路由权重:基于贡献对比的混合专家奖励模型的忠实响应级解释

Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast

Yifan Wang, Jinyi Mu, Mayank Jobanputra, Yu Wang, Soyoung Oh, Isabel Valera, Vera Demberg

arXiv 2608.06400首次发表:更新:

发表机构

Saarland University; University of California San Diego; Bielefeld University; Max Planck Institute for Software Systems; Max Planck Institute for Informatics(萨尔大学; 加利福尼亚大学圣迭戈分校; 比勒费尔德大学; 马克斯·普朗克软件系统研究所; 马克斯·普朗克信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对MoE奖励模型现有路由权重解释的局限性,本文提出CoCo响应级解释方法,经评估其解释质量更优且奖励建模准确率有竞争力,是该领域首个系统研究。

AI 中文摘要

奖励模型是从人类偏好中学习的核心,但识别其预测的驱动因素仍具挑战性。近期的稀疏混合专家(MoE)奖励模型通过将提示路由到专门的专家,并利用高路由权重的示例表征专家,以提升可解释性。然而,路由权重仅揭示专家“接收”哪些提示,而非其如何“评判”响应,仅能部分解释专家行为。为此,本文提出贡献对比(CoCo)响应级解释,利用贡献差异最大的选中-拒绝响应对忠实地表征专家角色,同时捕捉路由与偏好行为。经自动与人工评估,CoCo相比基于路由、基于得分及稀疏自编码器的替代方法,能产生更连贯、忠实且专业化的解释,同时保持具有竞争力的奖励建模准确率。据所知,这是首个针对MoE奖励模型解释方法的系统研究。

英文摘要

Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent sparse Mixture-of-Experts (MoE) reward models seek to improve interpretability by routing prompts to specialized experts and characterizing experts through examples with high routing weights. However, routing weights only reveal which prompts an expert $\textit{receives}$, not how it $\textit{judges}$ responses, providing only a partial account of expert behavior. We therefore propose $\textbf{Co}$ntribution-$\textbf{Co}$ntrast ($\textbf{CoCo}$) response-level interpretation, which faithfully characterizes experts' roles using chosen-rejected response pairs with the largest contribution contrasts, jointly capturing routing and preference behavior. Across automatic and human evaluations, CoCo yields more coherent, faithful, and specialized interpretations than router-based, score-based, and sparse autoencoder-based alternatives while maintaining competitive reward modeling accuracy. To the best of our knowledge, this is the first systematic study of interpretation methods for MoE reward models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑