任务空间模仿引导用于高效强化学习
Task-Space Imitation Guidance for Efficient Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
TIGER框架通过将模仿策略作为任务空间进度估计器并构建密集奖励,在稀疏奖励桌面操作任务中提升强化学习的样本效率与安全性。
中文摘要 AI 辅助
我们提出了任务空间模仿引导的高效强化学习框架(TIGER),这是一个针对稀疏奖励桌面机器人操作的奖励构建和预训练框架。TIGER将动作分块的模仿策略视为局部任务空间进度估计器,而非可执行控制器或动作先验:预测的动作块通过控制器感知的动作到运动映射转换为短视界末端执行器参考,强化学习代理接收朝向这些参考的密集进度奖励,同时稀疏环境奖励保持为主要目标。在预训练期间,TIGER使用模仿引导的前瞻信号来放松对预测能产生任务空间进展的动作的保守价值惩罚,减少早期在线强化学习中的流形外探索。在仿真和真实机器人实验中,TIGER提高了早期样本效率,减少了测量的安全违规,同时在评估任务上相对于先前的强化学习和模仿学习-强化学习基线匹配或提高了最终成功率。
英文摘要
We introduce Task-Space Imitation Guidance for Efficient Reinforcement Learning (TIGER), a reward-construction and pretraining framework for sparse-reward tabletop robotic manipulation. TIGER treats an action-chunked imitation policy not as an executable controller or action prior, but as a local task-space progress estimator: predicted action chunks are converted, using controller-aware action-to-motion mapping, into short-horizon end-effector references, and the RL agent receives dense progress rewards toward these references while the sparse environment reward remains the dominant objective. During pretraining, TIGER uses imitation-guided look-ahead signals to relax conservative value penalties for actions predicted to make task-space progress, reducing off-manifold exploration during early online RL. Across simulation and real-robot experiments, TIGER improves early sample efficiency and reduces measured safety violations while matching or improving final success rates relative to prior RL and IL-RL baselines on the evaluated tasks.
发表机构
- Sanctuary AI
机构由 AI 辅助整理,请以论文原文为准。