强化学习从人类反馈(RLHF)中的评分者状态偏差:一个审计框架
Rater State Bias in RLHF Preference Data: An Audit Framework
- Oles Honchar Dnipro National University(第聂伯罗国立奥列斯·冈察尔大学)
- Kyiv Institute of Modern Psychology and Psychotherapy(基辅现代心理与心理治疗研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究RLHF中评分者状态偏差这一结构化混淆,提出评分者状态变化是偏差来源,开发审计框架,定义相关概念,得出可证伪预测和效应大小阈值,提出审计协议和试点研究计划以分离偏差来源。
AI中文摘要:
我们在强化学习从人类反馈(RLHF)中识别出一种结构化混淆。成对偏好标签旨在反映比较的输出,但也可能反映注释期间评分者的状态。在持续的压力或痛苦条件下,评分者的偏好可能随时间变化。因此,偏好数据除了对响应质量的判断外,还能编码评分者状态。这些变化不同于普通分歧或随机标签噪声。它们依赖于状态,可在相似条件下工作的注释者之间共享,并能通过奖励建模和策略优化传播。我们提出评分者状态变化是RLHF偏好数据中结构化偏差的一个合理且可测试的来源。本文为研究这种偏差来源开发了一个假设和审计框架。我们定义了评分者状态变化、评分者状态混淆和相关评分者状态偏差。我们还使用词汇、语用、话语和安全相关特征将生存水平情感真实性定义为一种可测量的响应模式。我们分析了相关评分者状态偏差如何在聚合中幸存并进入学习到的奖励信号。我们得出了五个可证伪的预测和初始审计的效应大小阈值。最后,我们提出了一个可应用于公开可用的指令调整模型的审计协议和试点研究计划。我们不推断任何特定部署模型的训练历史。我们的目标是在RLHF偏好数据中分离出一个合理且可测试的结构化偏差来源。
英文摘要:
We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater's state during annotation. Under sustained stressful or distressing conditions, raters' preferences may shift over time, so that preference data encode rater state alongside judgments about response quality. We argue that, if present, such shifts would differ from ordinary disagreement or random label noise. They would be state dependent, could be shared across annotators under similar conditions, and would not necessarily cancel during aggregation, reward modeling, and policy optimization. We propose rater state shift as a plausible and testable source of structured bias in RLHF preference data. This paper develops a hypothesis and an audit framework for studying this source of bias. We define rater state shift, rater state confound, and correlated rater state bias. We also propose survival level emotional authenticity as a candidate output signature, defined by lexical, pragmatic, discourse, and safety features whose reliability and validity remain to be demonstrated. We show that systematic rater state bias can survive aggregation and may enter the learned reward signal. We state five testable predictions, together with effect size thresholds for an initial audit, and note which require proprietary data. Finally, we present an audit protocol and pilot study plan that can be applied to publicly available instruction tuned models. We do not infer the training history of any specific deployed model.