发表机构
University of Toronto(多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FOCUS是一种单阶段PPO框架,通过受控模态切换与表示对齐解决机器人操纵强化学习的训练-测试模态差距问题,在五项操纵任务中显著提升了测试成功率与学习效率。
AI 中文摘要
基于视觉的机器人操纵强化学习样本效率低,因为RGB-D观测维度高且存在噪声。仿真中可用的特权状态信息可加速训练,但测试时缺失该信息会造成训练-测试模态差距。我们提出FOCUS,这是一种单阶段PPO框架,其在特权状态上训练评论家,同时自动调节演员从RGB-D或特权状态潜变量收集rollout的行为。调节由两种模态诱导的动作分布之间的KL散度驱动,而表示对齐则鼓励两者间动作选择一致。这些机制共同限制了演员的RGB-D rollout,当RGB-D与特权状态潜变量的动作分布不一致时;当两者对齐时,RGB-D曝光度会增加,使在线策略训练向测试时使用的RGB-D输入偏移。在五项操纵任务中,相对于每个任务最强的测试时使用RGB-D的基线,FOCUS将平均测试成功率从0.71提升至0.93。当考虑每种方法的完整训练流程时,预算归一化的训练成功率AUC从0.47提升至0.65。在抓取放置任务中,测试成功率从0.47提升至0.86,而AUC从0.12提升至0.61,在固定交互预算下学习效率提升了5.0倍。
英文摘要
Vision-based reinforcement learning for robotic manipulation is sample-inefficient because RGB-D observations are high-dimensional and noisy. Privileged state information available in simulation can accelerate training, but its absence at test time creates a train-test modality gap. We propose FOCUS, a single-stage PPO framework that trains the critic on privileged state while automatically regulating whether the actor collects rollouts from RGB-D or privileged-state latents. Regulation is driven by the KL divergence between the action distributions induced by the two modalities, while representation alignment encourages consistent action selection across them. Together, these mechanisms limit RGB-D rollouts when the actor's action distributions from RGB-D and privileged-state latents disagree. As they align, RGB-D exposure increases, shifting on-policy training toward the RGB-D inputs used at test time. Across five manipulation tasks, FOCUS raises average test success from 0.71 to 0.93 relative to the strongest RGB-D-at-test baseline on each task. When accounting for each method's complete training pipeline, budget-normalized training-success AUC increases from 0.47 to 0.65. On Pick-and-Place, test success rises from 0.47 to 0.86, while AUC increases from 0.12 to 0.61, a 5.0x improvement in learning efficiency over the fixed interaction budget.
Comments38 pages, 9 figures