arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12140cs.ROcs.SYeess.SY

在强化学习中利用人在回路演示实现数字孪生驱动的机器人灵活性

Leveraging Human-In-The-Loop Demonstrations in Reinforcement Learning for Digital Twin-Driven Robot Flexibility

Yuzhu Sun, Mien Van, Nguyen Minh Nhat, Stephen McIlvanna, Sean McLoone

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出结合数字孪生、强化学习与人演示的人在回路在线训练框架,在Ufactory Xarm5机器人实验中,其双Actor框架在非最优演示下的成功率显著优于带模仿损失的对比方法。

中文摘要 AI 辅助

日益增长的自动化使协作机器人能在更多样化的环境中工作,这提升了对适应性的需求。我们提出一种结合数字孪生(DT)、强化学习(RL)与人演示的人在回路在线训练框架。与主要用于在任务执行前生成合成数据的数字孪生不同,我们的数字孪生通过相机馈送与物理系统实时同步,使虚拟机器人能从现实世界反馈中更新其观测值与策略。双Actor框架整合了模仿学习(IL),且未给RL Actor添加直接模仿损失,因此演示可引导适应性,而非手动重新编程。该框架在Ufactory Xarm5协作机器人上进行了验证,机器人末端执行器需在避开障碍物的同时到达目标位置。实验表明,该框架能在物理工作空间发生变化后恢复训练;在使用一组固定的非最优演示时,双Actor框架的最终成功率远高于两种给Actor添加模仿损失的方法。在通过虚拟现实(VR)收集的真实人类演示中也呈现相同规律:在演示从未达到目标的情况下,双Actor框架达到了83%-100%的平均确定性评估成功率,而两种带模仿损失的方法仅为0%-17%。

英文摘要

Growing automation makes collaborative robots work in more variable environments, increasing the need for adaptation. We propose a human-in-the-loop online training framework combining a digital twin (DT), reinforcement learning (RL), and human demonstrations. Unlike DTs used mainly to generate synthetic data before task execution, our DT is synchronized with the physical system in real time through camera feeds, allowing the virtual robot to update its observations and policy from real-world feedback. A dual actor framework integrates imitation learning (IL) without adding a direct imitation loss to the RL actor, so demonstrations can guide adaptation instead of manual reprogramming. The proposed framework is demonstrated on the Ufactory Xarm5 collaborative robot, where the robot's end-effector aims to reach the target position while avoiding obstacles. The experiments show that the framework can resume training after a change in the physical workspace and that, with a fixed set of non-optimal demonstrations, the dual actor framework achieves a much higher final success rate than two methods that add an imitation loss to the actor. The same pattern holds with real human demonstrations collected in virtual reality (VR): with demonstrations that never reach the goal, the dual actor framework reached 83-100% mean deterministic evaluation success, against 0-17% for the two imitation-loss methods.

发表机构

  • School of Electronics, Electrical Engineering and Computer Science, Queen’s University Belfast(贝尔法斯特女王大学电子、电气工程与计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑