零样本仿真到现实的接触丰富装配:基于本体感觉锚定的跨模态预训练
Proprioception-Anchored Cross-Modal Pretraining for Zero-Shot Sim-to-Real Contact-Rich Assembly
浏览论文内容
中文总结 AI 辅助
针对接触丰富装配的仿真到现实迁移难题,提出基于本体感觉锚定的跨模态预训练方法PACE,通过预测本体感觉状态转移监督视觉和力觉表征,实现零样本部署,平均成功率93.3%。
中文摘要 AI 辅助
接触丰富的装配任务因其需要亚毫米级空间精度以及在持续接触过程中对力的可靠解读而仍然具有挑战性。尽管基于仿真的强化学习提供了一种可扩展的训练范式,但视觉观测、接触动力学以及力/力矩(F/T)测量之间的差异常常限制策略的迁移。我们观察到,本体感觉在不同域间相对一致,因为校准后的关节位置和一致计算的关节速度在仿真与硬件之间紧密对齐。基于这一观察,我们提出了PACE(本体感觉锚定的跨模态编码器),它通过预测本体感觉状态转移来监督时间视觉和F/T表征。静态的域特定因素,包括光照、纹理和传感器偏差,包含关于关节运动的少量信息;因此,所提出的目标鼓励编码器抑制这些因素,同时保留任务相关的运动线索。在冻结的PACE特征上训练的策略在硬件上部署,无需现实世界微调或物体姿态跟踪。在四个接触丰富的装配任务中,PACE实现了93.3%的平均现实世界成功率,并且仅出现2.7个百分点的仿真到现实性能下降,同时对于显著降低基于姿态和基于学习融合基线的扰动保持鲁棒性。
英文摘要
Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained contact. Although simulation-based reinforcement learning offers a scalable training paradigm, discrepancies in visual observations, contact dynamics, and force/torque (F/T) measurements often limit policy transfer. We observe that proprioception is comparatively consistent across domains because joint positions are expressed in a shared calibrated coordinate system and joint velocities are computed consistently in simulation and on hardware. Based on this observation, we present PACE (Proprioception-Anchored Cross-Modal Encoder), which supervises temporal visual and F/T representations by predicting proprioceptive state transitions. Static domain-specific factors, including lighting, texture, and sensor bias, contain little information about joint motion; optimizing the proposed objective therefore suppresses their influence on the learned representation while retaining task-relevant motion cues. Policies trained on frozen PACE features are directly deployed on hardware without real-world fine-tuning or object-pose tracking. Across four contact-rich assembly tasks, PACE attains an average real-world success rate of 93.3\% and only a 2.7-percentage-point sim-to-real drop, while remaining robust to perturbations that substantially degrade pose-based and learned-fusion baselines.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。