arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

子空间推理实现高效的基于偏好的主动奖励学习

Subspace Inference Enables Efficient Active Reward Learning from Preferences

Yutai Zhou, Erdem Bıyık

arXiv 2609.04066首次发表:更新:

发表机构

University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PreferenceEKF方法,通过低维子空间的扩展卡尔曼滤波实现高效序贯贝叶斯推理,在D4RL等基准上较其他贝叶斯深度学习方法更优,助力RLHF的高效奖励学习。

AI 中文摘要

从人类反馈的强化学习(RLHF)是一种强大但样本效率低下的方法,用于从人类偏好中学习奖励模型,这使得主动学习成为合成信息丰富的偏好查询的关键组成部分。然而,对于大型神经网络奖励模型而言,主动学习所需的有效不确定性量化仍然是一个关键挑战。在本文中,我们提出了PreferenceEKF,这是一种样本高效的方法,通过将主动偏好学习构建为序贯贝叶斯滤波问题来跟踪奖励模型的不确定性。我们的方法不依赖于在完整神经网络参数空间上计算代价高昂的后验推理,而是在低维参数子空间内通过扩展卡尔曼滤波(extended Kalman filter)执行序贯推理,随着新偏好查询的到来不断更新奖励模型后验。我们的方法能够对神经网络参数进行可扩展采样,以高效计算主动奖励学习的获取函数。在D4RL和V-D4RL基准上的实验表明,与其他贝叶斯深度学习方法相比,我们的方法在样本效率、运行时间、可扩展性和校准方面表现更优,且学习到的奖励模型能产生具有竞争力的离线强化学习策略性能。这凸显了可扩展贝叶斯方法在RLHF中基于偏好的奖励建模方面的潜力。我们的代码可在此URL获取。

英文摘要

Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach for learning reward models from human preferences, making active learning a critical component in synthesizing informative preference queries. However, effective uncertainty quantification required for active learning remains a key challenge for large neural network reward models. In this paper, we introduce PreferenceEKF, a sample-efficient approach that tracks reward model uncertainty by framing active preference learning as a sequential Bayesian filtering problem. Instead of relying on computationally prohibitive posterior inference over the full neural network parameter space, our method performs sequential inference via an extended Kalman filter within a low-dimensional parameter subspace, continuously updating the reward model posterior as new preference queries arrive. Our approach enables scalable sampling of neural network parameters to efficiently compute acquisition functions for active reward learning. Experiments on the D4RL and V-D4RL benchmarks demonstrate that our approach achieves better sample efficiency, runtime, scalability, and calibration compared to other Bayesian deep learning approaches, and the learned reward models lead to competitive offline reinforcement learning policy performance. This highlights the potential of scalable Bayesian methods for preference-based reward modeling in RLHF. Our code is available at https://github.com/yutaizhou/bnn_pref.

CommentsPublished at TMLR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑