arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TaRL:从触觉演示中学习通用与物理奖励

TaRL: Learning General and Physical Rewards from Tactile Demonstrations

Po-Yi Wu, Dao-Jan Chang, Shang-Ya Hsiao, Hong-Ming Chen, Yu-Cheng Su, Tsung-Wei Ke

arXiv 2609.36785首次发表:更新:

发表机构

National Taiwan University; Delta Electronics(国立台湾大学; 台达电子)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TaRL 从触觉演示中学习奖励,用于接触丰富的操作任务,通过触觉变形图回归任务进度,提升样本效率和成功率,并跨物体泛化。

AI 中文摘要

接触丰富的操作任务要求机器人能够顺序地执行精确接触、保持稳定抓取并施加定向力。强化学习(RL)可以自动获取此类行为,但其性能依赖于奖励设计:稀疏奖励会降低学习效率,而稠密奖励难以指定。视觉奖励学习通过从无动作演示中推断奖励来解决这一问题。由于它仅基于视觉观测,因此无法捕获超出视觉目标的奖励。我们提出了触觉奖励学习(TaRL),这是一个从触觉演示中学习奖励的框架。TaRL 以一系列触觉变形图作为输入,并从成功和失败的演示中回归任务完成进度。由于 TaRL 捕获了局部机器人-物体交互,它提供了信息丰富的反馈以学习牢固抓取和正确方向的力;同时,它对场景布局的变化(如物体位置)具有鲁棒性。我们在四个仿真操作任务和两个真实世界任务上评估了 TaRL。作为塑形奖励使用时,它显著提高了样本效率和最终成功率,在仿真中将螺母螺纹连接的成功率从 34% 提升至 56%,在真实世界中将立方体抓取的成功率从 37% 提升至 97%。将触觉奖励与视觉奖励结合可进一步提高性能。TaRL 还能跨物体实例泛化:在盒子放置任务上训练,直接部署到罐子放置任务,显著提升了新任务上的策略学习。项目页面可在该 https URL 获取。

英文摘要

Contact-rich manipulation requires robots to sequence precise contacts, maintain stable grasps, and apply directed forces. Reinforcement learning (RL) can acquire such behaviors automatically, but its performance hinges on reward design: sparse rewards reduce the learning efficiency, while dense rewards are hard to specify. Visual reward learning addresses this by inferring rewards from action-free demonstrations. Because it conditions only on visual observations, it fails to capture rewards beyond visual goals. We propose Tactile Reward Learning (TaRL), a framework that learns rewards from tactile demonstrations. TaRL takes a sequence of tactile deformation maps as input, and regresses task-completion progress from both successful and failed demonstrations. Because TaRL captures local robot-object interaction, it provides informative feedback to learn firm grasps and correctly directed forces; meanwhile, it is robust to changes in scene layout such as object position. We evaluate TaRL on four manipulation tasks in simulation and two in the real world. Used as a shaping reward, it substantially improves both sample efficiency and final success rate, raising success on Nut threading from 34% to 56% in simulation and on cube pickup from 37% to 97% in the real world. Combining tactile with visual rewards improves performance further. TaRL also generalizes across object instances: trained on box placement and directly deployed to can placement, it significantly improves policy learning on the new task. Project page is available at https://embodiedai-ntu.github.io/tarl.

Comments7 pages, 13 figures. Project page: https://embodiedai-ntu.github.io/tarl

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑