BiasReducer:奖励模型的自适应偏差缓解
BiasReducer: Adaptive Bias Mitigation for Reward Models
AI总结:
提出BiasReducer轻量级框架,通过编辑线性奖励头并自适应选择相关编辑,缓解奖励模型对表面属性的偏差,提升鲁棒性并在多个基准上显著优于训练基线。
AI中文摘要:
奖励模型对大型语言模型(LLM)生成的回复进行评分,并引导LLM的训练朝着符合人类偏好的方向进行。然而,奖励模型可能偏好诸如长度或置信度等表面属性,导致LLM生成得分更高但并非更正确的回复。现有的缓解方法要么重新训练奖励模型,要么对已知偏差(如偏好较长回复)应用固定修正。重新训练需要额外的数据和计算资源,而现有的编辑方法需要预先指定目标偏差,并对该偏差使用固定编辑。为此,我们提出了BiasReducer,一个轻量级框架,仅编辑线性奖励头,并为每个新数据集选择相关的编辑。首先,BiasReducer使用稀疏自编码器(SAE)风格的编码器来学习奖励模型对哪些属性(如长度和置信度)敏感。其次,它通过确定调整奖励头的方向和调整幅度,学习如何减少奖励模型对每个属性的依赖。第三,对于新数据集,它根据属性对奖励分数的影响进行排序,选择相关属性,并相应编辑奖励模型。BiasReducer持续提高了奖励模型对表面回复属性偏差的鲁棒性。在五个奖励模型上,BiasReducer-M在三个基准上平均提高了8.3、18.0和6.9个百分点,优于两个基于训练的基线。这些增益可迁移到下游任务,减少了不必要的冗长和谄媚,同时保持了可比的评判质量。
英文摘要:
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing editing methods require the target bias to be specified in advance and use a fixed edit for that bias. To this end, we propose BiasReducer, a lightweight framework that edits only the linear reward head and selects the relevant edits for each new dataset. First, BiasReducer uses a sparse autoencoder (SAE)-style encoder to learn which attributes (e.g., length and confidence) the reward model is sensitive to. Second, it learns how to reduce the reward model's dependence on each attribute by determining which direction to adjust the reward head and how much to adjust it. Third, for a new dataset, it ranks the attributes by their influence on reward scores, selects the relevant ones, and edits the reward model accordingly. BiasReducer consistently improves reward-model robustness to biases toward superficial response attributes. Across five reward models, BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming the two training-based baselines. The gains transfer downstream, reducing unnecessary verbosity and sycophancy while maintaining comparable judged quality.