PHIRL:将学习奖励与任务进度对齐以进行逆强化学习
PHIRL: Aligning Learned Rewards with Task Progress for Inverse Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
PHIRL提出一种联合演示与进度反馈的逆强化学习框架,通过四维对齐学习稳健奖励,在少量标注下显著提升机器人任务回报与成功率。
中文摘要 AI 辅助
人类演示提供了密集的策略级信息,但有时缺乏局部精度。人类反馈提供了准确的局部批评,但提供的是稀疏的评估而非直接的策略指导。我们提出了进度启发式逆强化学习(PHIRL),一种数据高效的框架,通过联合利用演示和反馈来学习稳健的奖励函数。具体来说,我们使用进度,一种描述累积任务完成情况的反馈模态。PHIRL通过逆强化学习从演示中迭代推断奖励函数,计算学习到的奖励在带有进度标注的演示上的值,并将奖励与进度标注在四个维度上对齐。我们在真实和模拟机器人任务上评估PHIRL,并额外使用微调的视觉语言模型来提供进度反馈。结果表明,PHIRL显著优于基线,在仅有百分之二十的演示被标注的情况下,实现了更高的环境回报奖励和任务成功率。对奖励黑客场景的分析表明,PHIRL学习到的奖励函数在对抗利用方面是可靠的。
英文摘要
Human demonstrations provide dense policy-level information but sometimes lack local precision. Human feedback presents accurate local critiques, but offers sparse evaluations rather than direct policy guidance. We propose Progress-Heuristicized Inverse Reinforcement Learning (PHIRL), a data-efficient framework that learns robust reward functions by jointly leveraging demonstrations and feedback. Specifically, we use progress, a feedback modality that describes cumulative task completion. PHIRL iteratively infers a reward function from demonstrations via inverse reinforcement learning, calculates the learned rewards over the progress-annotated demonstrations, and aligns the rewards with progress annotations over four dimensions. We evaluate PHIRL on real and simulated robot tasks, with additional exploration using a fine-tuned vision-language model to provide progress feedback. Results demonstrate that PHIRL significantly outperforms the baselines, achieving substantially higher environmental return rewards and task success with only twenty percent of demonstrations annotated. Analysis of reward-hacking scenarios demonstrates that PHIRL learned reward functions are reliable against exploitation.