PredTac:利用预测触觉学习接触丰富的操作
PredTac: Learning Contact-Rich Manipulation with Predicted Touch
查看机构详情
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
PredTac通过视觉和机器人状态预测触觉,替代物理触觉传感器,在仿真和真实机器人上提升接触丰富操作成功率,接近真实触觉性能。
中文摘要 AI 辅助
接触丰富的操作受益于触觉反馈,然而物理触觉传感器引入了硬件、校准、同步和维护成本,这些成本使策略学习和部署变得复杂。我们将预测触觉表述为测量触觉输入的替代方案,并提出PredTac,这是一个从因果视觉观察和机器人状态中学习推断触觉状态,并使用预测触觉作为策略学习和执行的显式接口的框架。首先,使用触觉监督训练一个触觉预测器,然后在后续策略训练或执行过程中,无需测量触觉输入即可提供接触信息。我们在仿真和真实机器人上评估了PredTac在三个接触丰富操作任务中的表现,并进一步考察了策略性能如何依赖于预测接触内容。在仿真目标偏移评估中,预测触觉策略在USB、Barbed-spike和Valve任务上的成功率分别为27.0%、52.0%和44.7%,比视觉基线提高了8.0至13.7个百分点。在真实机器人上,预测触觉ACT在USB插入、Barbed拔取和Valve旋转上的成功率分别为70.0%、50.0%和90.0%,三任务平均成功率为70.0%,接近测量触觉ACT的72.2%,并显著优于视觉ACT的21.1%。固定策略干预进一步表明,性能对预测接触的空间结构敏感,在固定值分布下进行空间重排使Valve成功率降低了10.7个百分点。这些结果表明,预测触觉可以在无需将触觉感知作为策略输入的情况下,为接触丰富的操作提供有用的接触信息。
英文摘要
Contact-rich manipulation benefits from tactile feedback, yet physical tactile sensors introduce hardware, calibration, synchronization, and maintenance costs that complicate policy learning and deployment. We formulate predicted touch as an alternative to measured tactile input and present PredTac, a framework that learns to infer tactile states from causal visual observations and robot states and uses the predicted touch as an explicit interface for policy learning and execution. A tactile predictor is first trained with tactile supervision and then used to provide contact information without requiring measured tactile input during downstream policy training or execution. We evaluate PredTac across three contact-rich manipulation tasks in simulation and on a real robot, and further examine how policy performance depends on the predicted contact content. In simulation goal-offset evaluations, predicted-touch policies achieve 27.0%, 52.0%, and 44.7% success on USB, Barbed-spike, and Valve, respectively, improving over the visual baseline by 8.0-13.7 percentage points. On the real robot, predicted-touch ACT achieves 70.0%, 50.0%, and 90.0% success on USB insertion, Barbed extraction, and Valve rotation, respectively, with a three-task mean of 70.0%, approaching measured-touch ACT at 72.2% and substantially outperforming visual ACT at 21.1%. Fixed-policy interventions further show that performance is sensitive to the spatial structure of predicted contact, with spatial rearrangement at fixed value distributions reducing Valve success by 10.7 percentage points. These results demonstrate that predicted touch can provide useful contact information for contact-rich manipulation without requiring tactile sensing as a policy input.