AI 中文总结
本研究提出GORDON框架,通过无动作视频演示学习以对象为中心的密集奖励,自动分解长程操作任务,在7个基准任务上平均成功率达74.4%,显著优于现有基线。
AI 中文摘要
利用强化学习学习长程操作技能仍面临诸多挑战,包括奖励设计的复杂性、稀疏奖励的引导有限以及手动子任务标注成本高昂。视觉演示可为奖励学习提供监督,但从原始像素学习的奖励易受视觉变化、背景外观和机器人运动的影响而变得脆弱。本研究提出GORDON,一种基于图的以对象为中心的奖励学习框架,可从无动作的视频演示中学习密集奖励。每个视觉场景被表示为检测到的对象和空间关系构成的图,通过自监督方式训练图神经网络,将这些图嵌入到与任务对齐的潜在空间中。为使表示与语义任务进度对齐,引入了活动感知加权池化机制,该机制会强调与任务相关的对象,同时屏蔽机器人主导的运动。随后,将当前状态在学习到的潜在空间中与演示目标配置的距离计算为密集奖励,以此作为任务进度的度量。在长程任务中,该奖励的时间分布会呈现阶段性的对象状态转换,无需手动分段即可实现自动子任务发现。发现的分段随后用于训练特定子任务的奖励和可顺序组合的专用策略。在MAGICAL和ManiSkill3基准的7个操作任务上进行的实验表明,本研究的以对象为中心的奖励在短程设置中提升了强化学习效果,并通过自动分解实现了复杂长程任务中的成功策略学习,在长程任务上的平均成功率达到74.4%(相比最佳学习基准平均提升约35个百分点,相比oracle平均提升约25个百分点)。
英文摘要
Learning long-horizon manipulation skills with reinforcement learning remains challenging due to the complexity of reward design, the limited guidance of sparse rewards, and the high cost of manual subtask annotation. Visual demonstrations can provide supervision for reward learning, but rewards learned from raw pixels can be brittle and sensitive to visual variation, background appearance, and robot motion. In this work, we propose GORDON, a graph-based object-centric reward learning framework that learns dense rewards from action-free video demonstrations. Each visual scene is represented as a graph of detected objects and spatial relations, and a graph neural network is trained in a self-supervised manner to embed these graphs into a task-aligned latent space. To align the representation with semantic task progress, we introduce an activity-aware weighted pooling mechanism that emphasizes task-relevant objects while masking robot-dominated motion. The dense reward is then computed as distances in the learned latent space of the current state to demonstrated goal configurations, providing a measure of task progress. In long-horizon tasks, the temporal profile of this reward reveals stage-wise object-state transitions, enabling automatic subtask discovery without manual segmentation. The discovered segments are then used to train subtask-specific rewards and specialized policies that are composed sequentially. Experiments on seven manipulation tasks on MAGICAL and ManiSkill3 benchmarks show that our object-centric reward improves reinforcement learning in short-horizon settings and enables successful policy learning in complex long-horizon tasks through automatic decomposition, achieving an average success rate of 74.4% across the long-horizon tasks (on average approximately +35 p.p. vs. best learned baseline and approximately +25 p.p. vs. oracle).
Comments9 pages, 6 figures, preprint. Project page: https://andreaprotopapa.github.io/graph-reward-learning/