从混合质量部署经验中学习机器人操作
Learning from Mixed-Quality Deployment Experience for Robot Manipulation
浏览论文内容
中文总结 AI 辅助
针对机器人部署后混合质量经验利用问题,提出预测性动作块学习(PACL),通过块级评论家与未来预测增强价值估计,并指导扩散演员学习,实验证明其优于现有基线。
中文摘要 AI 辅助
在真实环境中部署的机器人策略自然积累混合质量的体验,包括成功执行、部分进展和失败。尽管这些轨迹为进一步学习提供了宝贵信息,但直接将它们纳入模仿学习可能会强化不良行为,而离线强化学习在稀疏奖励和有限数据覆盖下往往遭受不可靠的价值估计。我们考虑一个实际的部署后设置,其中学习仅依赖于自然积累的自主轨迹,无需额外的人工纠正或探索性交互。为了有效利用此类经验,我们提出了预测性动作块学习(PACL)。PACL首先学习一个预测性的块级评论家,该评论家评估时间上扩展的动作序列,并通过未来潜在预测增强时间差分学习,为长视野价值估计提供更丰富的监督。然后,学习到的评论家将块级Q值转换为离散的质量条件,指导扩散演员从这些混合质量经验中联合学习,而不将所有行为视为等效监督。在推理时,演员生成多个动作块,评论家选择价值最高的候选。在模拟和真实世界机器人操作任务上的实验表明,PACL持续改进预训练策略,并优于强模仿学习和离线强化学习基线。
英文摘要
Robot policies deployed in real environments naturally accumulate mixed-quality experience, including successful executions, partial progress, and failures. Although these rollouts provide valuable information for further learning, directly incorporating them into imitation learning may reinforce undesirable behaviors, while offline reinforcement learning often suffers from unreliable value estimation under sparse rewards and limited data coverage. We consider a practical post-deployment setting where learning relies only on naturally accumulated autonomous rollouts, without additional human corrections or exploratory interaction. To effectively exploit such experience, we propose Predictive Action Chunk Learning (PACL). PACL first learns a predictive chunk-level critic that evaluates temporally extended action sequences and augments temporal difference learning with future latent prediction, providing richer supervision for long-horizon value estimation. The learned critic then converts chunk-level Q-values into discrete quality conditions, which guide a diffusion actor to learn jointly from these mixed-quality experiences without treating all behaviors as equivalent supervision. At inference, the actor generates multiple action chunks and the critic selects the highest valued candidate. Experiments across simulated and real-world robot manipulation tasks show that PACL consistently improves the pretrained policy and outperforms strong imitation learning and offline reinforcement learning baselines.
发表机构
- Nanyang Technological University(南洋理工大学)
- Chongqing Changan Automobile Co., Ltd(重庆长安汽车股份有限公司)
- Guangzhou Cloudbutterfly Technology Co., Ltd.(广州云蝶科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。