反思性语言奖励设计用于多元对齐
Reflective Verbal Reward Design for Pluralistic Alignment
- University of Waterloo(滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种通过反思性对话构建个性化语言奖励模型的方法,以解决人类价值观多元性下聚合奖励模型压制少数偏好问题,实验显示准确性提升9-12%且样本效率更高。
AI中文摘要:
AI智能体通常通过基于人类反馈的强化学习(RLHF)与“人类价值观”对齐,其中从聚合的人类反馈中学习一个单一的奖励模型,并用于对齐智能体的行为。然而,人类价值观并非同质的——不同的人持有不同甚至相互冲突的价值观。将反馈聚合到单一奖励模型中有可能不成比例地压制少数群体的偏好。为了解决这个问题,我们提出了一种新颖的奖励建模方法,用于学习个性化的奖励模型。我们的方法使用一个语言模型引导用户进行反思性对话,在对话中用户批评智能体行为并构建自己的偏好。这个包含用户反思和批评示例的个性化对话历史,随后被用作另一个语言模型的上下文,该语言模型充当个性化奖励函数(我们称之为“语言奖励模型”),用于评估新的轨迹。在30名参与者的研究中,我们的方法在准确性上比非反思性语言奖励模型提高了9-12%,同时比传统监督学习方法更具样本效率。
英文摘要:
AI agents are commonly aligned with "human values" through reinforcement learning from human feedback (RLHF), where a single reward model is learned from aggregated human feedback and used to align an agent's behavior. However, human values are not homogeneous--different people hold distinct and sometimes conflicting values. Aggregating feedback into a single reward model risks disproportionately suppressing minority preferences. To address this, we present a novel reward modeling approach for learning individualized reward models. Our approach uses a language model to guide users through reflective dialogues where they critique agent behavior and construct their preferences. This personalized dialogue history, containing the user's reflections and critiqued examples, is then used as context for another language model that serves as an individualized reward function (what we call a "verbal reward model") for evaluating new trajectories. In studies with 30 participants, our method achieved a 9-12% improvement in accuracy over non-reflective verbal reward models while being more sample efficient than traditional supervised learning methods.