arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02951cs.LGcs.AI

SP3O:无需奖励建模的分段偏好强化学习

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

Evan Assmus, Qining Zhang, Lei Ying

首次发表
浏览论文内容

中文总结 AI 辅助

SP3O是一种无需奖励建模的新型PbRL算法,利用分段级偏好反馈,通过PPO型损失函数优化策略,在机器人控制和LLM微调等长视界任务中性能优于现有算法。

中文摘要 AI 辅助

针对一般随机马尔可夫决策过程(MDPs)的偏好强化学习(PbRL)通常需要训练奖励模型。现有无奖励模型的方法要么局限于多臂老虎机或确定性MDPs,如DPO或P3O,要么使用零阶、无梯度优化,其收敛速度通常慢于基于梯度的算法。此外,现有无奖励模型的偏好强化学习算法几乎仅使用轨迹级反馈,当轨迹较长时,人类评估者需付出大量精力。而分段更短,更易于比较和评估。本文提出一种新颖的无奖励模型、无评论家、基于梯度的PbRL算法,适用于分段偏好,命名为分段近端策略优化(SP3O)。SP3O利用分段级偏好反馈,通过离策略重要性采样构建准确的策略值差异估计器,再通过PPO型损失函数计算策略梯度。本文为该算法提供理论基础,分析分段长度选择的权衡,并在机器人控制和大语言模型(LLM)微调场景中与其他PbRL/RLHF算法对比实验,证明其性能提升,尤其在长视界任务中表现突出。

英文摘要

Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. Existing reward-model-free methods are either restricted to bandits or deterministic MDPs, such as DPO or P3O, or use zeroth-order, gradient-free optimization, which in general exhibits a slower convergence rate than gradient-based algorithms. Furthermore, existing reward-model-free preference-based RL algorithms almost exclusively use trajectory-level feedback, which can require significant effort from a human evaluator when trajectories are long. On the other hand, segments are much shorter, so they are easier to compare and evaluate. In this paper, we introduce a novel reward-model-free, critic-free, and gradient-based PbRL algorithm compatible with segment preferences named Segment Pairwise Proximal Policy Optimization (SP3O). SP3O utilizes segment-level preference feedback to construct an accurate policy value difference estimator via off-policy importance sampling, and then uses the estimator to compute the policy gradient via a PPO-type loss function. We provide a theoretical basis for the algorithm and analyze the tradeoff in choosing the segment length. We also evaluate it experimentally against other PbRL/RLHF algorithms in robotic control and LLM finetuning settings to show its improved performance, especially in long-horizon tasks.

发表机构

  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

↑