arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33989cs.CLcs.LG

RewardExplainer:从反事实偏好反馈中学习奖励模型解释

RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback

Jingyi He, Nier Wu, Shuang Liu, Xin Wang, Mengnan Du, Xia Hu

首次发表
浏览论文内容

中文总结 AI 辅助

提出RewardExplainer框架,通过反事实反馈训练可复用的奖励模型解释器,生成可干预的自然语言评分机制,提升解释忠实度并用于识别偏见和去偏,增强模型鲁棒性。

中文摘要 AI 辅助

奖励模型(RMs)是大语言模型后训练中的关键组成部分,为后续的强化学习提供奖励信号。然而,传统的判别式奖励模型通常只输出标量分数,这使得难以识别与其评分决策相关的响应行为。现有的解释方法往往依赖于预定义的高层属性,并且需要对每个响应对进行重复的反事实干预来验证候选解释,缺乏一种利用奖励模型的反馈来训练可复用解释器的闭环机制。为解决这一问题,我们提出了RewardExplainer,一个通过反事实改写从目标奖励模型获取反馈,并利用该反馈进一步优化解释器的框架。RewardExplainer生成开放式的、原子性的、可干预的自然语言评分机制,使解释更加具体、可读和可操作。它进一步将反事实反馈转化为偏好监督,使解释器能够比单次生成更忠实地捕捉目标奖励模型的评分偏好和敏感行为。在多个目标奖励模型和解释器骨干上的大量实验显示出一致的改进。除了解释之外,我们利用生成的机制来识别潜在的偏见模式,并构建有针对性的去偏数据用于微调奖励模型,从而提高在奖励黑客基准上的鲁棒性。

英文摘要

Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Existing interpretation methods often rely on predefined high-level attributes and require repeated counterfactual interventions for each response pair to validate candidate explanations, lacking a closed-loop mechanism that uses RMs' feedback to train a reusable explainer. To address this, we propose RewardExplainer, a framework that obtains feedback from the target reward model through counterfactual rewriting and uses this feedback to further optimize the explainer. RewardExplainer generates open-ended, atomic, and intervenable natural-language scoring mechanisms, making explanations more concrete, readable, and actionable. It further converts counterfactual feedback into preference supervision, enabling the explainer to more faithfully capture the target RM's scoring preferences and sensitive behaviors than single-pass generation. Extensive experiments across multiple target RMs and explainer backbones show consistent improvements. Beyond interpretation, we use the generated mechanisms to identify potential bias patterns and construct targeted debiasing data for fine-tuning the reward model, improving robustness on reward-hacking benchmarks.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Shanghai Jiao Tong University(上海交通大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Jilin University(吉林大学)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑