发表机构
The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对高质量遥操作数据集收集成本高的问题,提出通过逆向简单任务进行硬策略学习的框架,含闭环数据收集、分层数据细化和迭代策略学习方法,经实验验证该方法能高效稳定地训练复杂操作任务,提升硬任务成功率。
AI 中文摘要
高质量的遥操作数据集收集成本高昂,尤其是对于困难任务。许多任务存在方向不对称性,逆向简单任务轨迹可作为硬任务的可扩展监督信号,降低手动演示收集成本。但逆向数据可能有噪声,直接训练可能产生次优策略。为此提出一个硬策略学习框架,包括闭环数据收集管道、分层数据细化管道和迭代策略学习方法。通过结合自动化收集、分层细化和迭代学习,该方法能对复杂高精度操作任务进行可扩展、可靠的训练。在两个模拟基准和真实机器人实验中,该方法比基于逆向和强化学习的基线提高了硬任务成功率,数据效率更高,训练更稳定,且无需大量硬任务遥操作。
英文摘要
High-quality teleoperation datasets are costly to collect, particularly for hard tasks. We observe that many tasks exhibit directional asymmetry: completing the forward hard task is difficult, whereas reversing it by relaxing or disrupting the environment is comparatively easy. This suggests that reversed easy-task trajectories can serve as a scalable supervision signal for the hard task, reducing the cost of manual demonstration collection. However, reversed data can be noisy, and directly training on it may yield suboptimal policies. To enable largely automated acquisition and effective use of reversed data, we propose a teleoperation-cost effective framework for hard policy learning via temporal reversal of easy tasks, consisting of three key components: a closed-loop data collection pipeline that alternates between hard-task and easy-task policies to autonomously reset the environment and generate diverse trajectories; a hierarchical data refinement pipeline that temporally inverts easy-task rollouts and filters low-quality motion using kinematic priors and a critic-guided advantage filter; and an iterative policy learning method that trains the hard-task policy using both initial reversed easy-task demonstrations and the filtered reversed data in a continuous online learning loop. By combining automated collection, hierarchical refinement, and iterative learning, our method enables scalable, reliable training of complex, high-precision manipulation tasks. Across two simulated benchmarks and real-robot experiments, we demonstrate that our method improves hard-task success rates with higher data efficiency and more stable training compared to reversal-based and reinforcement-learning baselines, without requiring extensive hard-task teleoperation.
Comments17 pages, 17 figures