arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24484cs.LGcs.CL

奖励模型记住了什么?

What do Reward Models Memorize?

Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova

首次发表
浏览论文内容

中文总结 AI 辅助

研究通过在两个人类偏好数据集上测量反事实记忆,探究判别训练的奖励模型记住了什么。发现奖励模型存在记忆分配错误、记住特定数据集捷径、过度概括简单启发式关联等问题,表明此类模型有偏差,无法在上下文相关场景中判断响应质量。

中文摘要 AI 辅助

本文通过在两个人类偏好数据集上测量反事实记忆,研究了经过判别训练的奖励模型(RMs)记住了什么。我们发现,奖励模型存在如下问题:一是将记忆错误分配到容易、有高边际的偏好对;二是记住特定数据集的捷径(如模型标识、用户采样策略);三是面对未见的偏好对时过度概括人类偏好的简单启发式关联(如长度、合规性)。总体而言,我们的发现表明,从人类偏好数据中对奖励模型进行判别训练会导致有偏差的奖励模型,它们尚无法在上下文相关场景中判断响应质量。

英文摘要

This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.

补充信息

↑