发表机构
Resolution(睿策尔(Resolution))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文将基于奇异学习理论的模式化方法用于去偏奖励模型,利用偏好对的敏感性重新加权,在RM-Bench Hard上取得显著提升,权重可跨模型迁移,为奖励模型去偏提供了新方案。
AI 中文摘要
已知基于人类偏好训练的奖励模型会存在长度、格式及其他风格偏差。本文采用模式化方法,该方法根据每个偏好对在基准损失后验期望值上的测量效应(即其敏感性)对其进行重新加权,以去偏在Skywork-Reward-Preference v0.2上训练的Gemma 2 9B Instruct奖励模型。我们在RM-Bench Hard上取得了+14.2±1.2个百分点的提升(该划分中风格线索与正确性相悖,结果为5次随机种子的均值±标准误),同时保持了整体RM-Bench准确率,与已发表的最接近对比方法SteerRM所报告的Hard划分最大提升(+13.2个百分点)相当。我们在一个简单案例中证明,该重新加权可通过将干预的副作用(对RM-Bench安全子集的回归)追溯到一小类训练对来解释,且我们通过 ablation 验证了这一点。这些权重还具有可迁移性:在Gemma 2 9B上计算得到的权重无需重新计算即可去偏Gemma 2 2B和27B,并可部分迁移到Llama 3.1 8B。这是模式化(一种基于奇异学习理论的方法)首次应用于小型模型和合成任务之外的场景。
英文摘要
Reward models trained on human preferences are known to suffer from length, formatting, and other stylistic biases. In this paper we use patterning, which reweights each preference pair according to its measured effect on posterior expectation values of benchmark losses (its susceptibility), to debias a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preference v0.2. We obtain $+14.2 \pm 1.2$ pp on RM-Bench Hard, the split where style cues point against correctness (mean $\pm$ s.e.\ over 5 seeds), with overall RM-Bench accuracy preserved, comparable to the strongest Hard-split gain reported by the closest published comparator (SteerRM, $+13.2$ pp). We demonstrate in a simple case that the reweighting is interpretable by tracing a side effect of the intervention (a regression on a safety subset of RM-Bench) to a small class of training pairs, which we confirm by ablation. The weights also transfer: those computed on Gemma 2 9B debias Gemma 2 2B and 27B with no recomputation, and transfer partially to Llama 3.1 8B. This is the first application of patterning, a program grounded in singular learning theory, beyond small models and synthetic tasks.