arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEMUR:基于偏好反馈的多目标强化学习对齐学习

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

Manith Adikari, Bei Peng, Samuele Vinanzi, Angelo Cangelosi

arXiv 2607.29559首次发表:更新:

AI 中文总结

本研究提出LEMUR框架,通过联合学习策略与多目标奖励模型,从人类偏好反馈中学习平衡多目标的最优策略,其在基准多目标任务上性能优于基线方法。

AI 中文摘要

强化学习(RL)系统通常使用单一、明确的标量奖励函数进行训练,但现实决策任务往往涉及多个相互冲突的目标,如性能与效率的权衡,此时真实奖励函数难以指定或获取。多目标强化学习(MORL)通过将奖励建模为向量来解决此类权衡问题,但现有方法通常假设每个目标都有明确的奖励函数,继承了单目标RL面临的相同挑战。同时,基于偏好的强化学习(PbRL)通过从人类反馈中学习奖励,在无需预定义奖励函数的复杂任务中展现出巨大潜力,但大多在单目标场景中研究。本研究提出LEMUR:基于偏好反馈的多目标强化学习对齐学习,这是一种智能体通过与多名人类的偏好交互学习最优多目标策略的新框架。该方法从人类反馈中联合学习策略和多个特定目标的奖励模型,使智能体在学习过程中有效平衡相互冲突的目标。在多种基准多目标任务上对LEMUR进行评估,实验结果表明其性能优于基线方法,为解决无需预定义奖励函数的多目标决策任务提供了有前景的方向。

英文摘要

Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, competing objectives, such as performance versus efficiency, where ground-truth reward functions are difficult to specify or inaccessible. While Multi-Objective RL (MORL) addresses such trade-offs by modeling rewards as vectors, existing approaches typically assume access to a well-specified reward function for each objective, inheriting the same challenges faced by single-objective RL. Meanwhile, Preference-based RL (PbRL) has shown great potential in solving complex tasks without access to a pre-defined reward function through reward learning from human feedback, yet has largely been studied in single-objective settings. In this work, we bridge this gap with LEMUR: Learning to Align with Multi-Objective Reinforcement Learning with Preference feedback, a novel framework where an agent interactively learns from the preferences of multiple humans to learn optimal multi-objective policies. Our approach jointly learns policies and multiple objective-specific reward models from human feedback, enabling agents to effectively balance competing objectives during learning. We evaluate LEMUR on a variety of benchmark multi-objective tasks, and empirical results demonstrate its superior performance over baseline methods. Our method presents a promising direction for solving multi-objective decision-making tasks without pre-defined reward functions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑