arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25350cs.LGcs.RO

超越成对反馈:用于基于偏好的奖励学习的列表式视觉-语言监督

Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning

Srivalli Katkuri, Maxwell Kawada, Juan Wachs

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出首个结合视觉-语言模型生成偏好与普拉科特-卢斯模型的列表式奖励学习框架,在Meta-World操纵任务中,其表现与基线相当且更灵活,最佳配置达86%平均成功率。

中文摘要 AI 辅助

视觉-语言模型(VLMs)已成为强化学习中强大的监督源,使智能体在训练过程中能利用丰富的语义知识。受人类反馈强化学习(RLHF)中基于偏好的奖励学习(PbRL)成功的启发,视觉-语言模型生成的基于图像的偏好为学习奖励函数提供了有效来源,这可通过布拉德利-特里(BT)模型对两个结果进行视觉比较来实现。然而,这种成对表述每次仅利用两个观测值,尽管VLMs能够对多个候选进行排序。普拉科特-卢斯(PL)表述可利用列表式排名而非成对偏好来构建奖励模型,从而更适配基于VLM的排序。在本研究中,据我们所知,我们推出了首个将VLM生成的偏好与普拉科特-卢斯模型结合用于奖励学习的框架。我们在Meta-World操纵任务上评估了该方法,结果显示,普拉科特-卢斯(PL)奖励模型从VLM生成的排名中训练机器人策略的效果,与成对布拉德利-特里、K-wise布拉德利-特里和RL-VLM-F基线相当。在所有环境中,至少有一种PL排名规模(K∈{3,4,5})在平均成功率上始终表现与其他方法相当或更优。与限于K=2的成对方法不同,PL支持不同的排名规模,因此可适配环境和所需的反馈格式。我们的最佳PL配置达到86%的平均最终成功率,且在Drawer Open任务上与Oracle基线表现相当。总体而言,这些结果表明,列表式VLM偏好监督是强化学习奖励学习中一种具有竞争力且灵活的方法。

英文摘要

Vision-language models (VLMs) have emerged as a powerful source of supervision for reinforcement learning, enabling agents to leverage rich semantic knowledge during training. Inspired by the success of preference-based reward learning (PbRL) in reinforcement learning from human feedback (RLHF), vision-language model generated image-based preferences provide an effective source for learning reward functions. This can be done by visually comparing two outcomes through the Bradley-Terry (BT) model. However, this pairwise formulation utilizes only two observations at a time, despite VLMs being capable of ranking multiple candidates. The Plackett-Luce (PL) formulation can shape a reward model with listwise rankings as opposed to pairwise preferences, allowing for a more suited use of a VLM based ranking. In this work, to our knowledge, we introduce the first framework that combines VLM-generated preferences with the Plackett-Luce model for reward learning. We evaluate our approach on Meta-World manipulation tasks and show that Plackett-Luce (PL) reward models can train robotic policies from VLM-generated rankings as effectively as pairwise Bradley-Terry, $K$-wise Bradley-Terry, and RL-VLM-F baselines. Across all environments, at least one PL ranking size ($K \in \{3,4,5\}$) consistently performs with or outperforms other methods in mean success rate. Unlike pairwise methods, which are restricted to $K=2$, PL supports different ranking sizes and can therefore be adapted to the environment and desired feedback format. Our best PL configuration achieves an 86% mean final success rate and matches the Oracle baseline on Drawer Open. Overall, these results demonstrate that listwise VLM preference supervision is a competitive and flexible approach to reward learning for reinforcement learning.

补充信息

↑