arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PSR:面向接触丰富操作任务的预测性感觉运动表征学习

PSR: Predictive Sensorimotor Representation Learning for Contact-Rich Manipulation

Shengbao Li, Peng Xu, Chao Tang, Hao Wei, Jiaheng Wang, Hong Yin, Jiangtao Chen, Jinxuan Zhu, Zhong Zhou, Mengfan Wang, Tingguang Li

arXiv 2609.21753首次发表:更新:

发表机构

Samsung R&D Institute China-Beijing(三星电子中国研究院北京)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PSR框架,通过多模态Transformer预训练学习预测性表征层次,增强视觉运动策略,在接触丰富操作中实现91.7%成功率,显著优于基线。

AI 中文摘要

接触丰富的操作任务要求策略能够超越视觉观察,通过对接触力、机器人构型和交互历史进行推理来生成精确动作。现有方法被动地依赖力反馈,而非主动预测未来的接触动力学,这限制了其生成高精度动作的能力。为解决这一问题,我们提出了预测性感觉运动表征(PSR)学习框架,该框架从多模态感觉运动信号中学习预测性表征的层次结构,并将其整合到视觉运动策略的动作流中。具体而言,在预训练阶段,一个多模态Transformer通过联合预测未来的交互动力学来学习预测性表征的层次结构。学习到的层次结构随后增强动作流,使所得策略能够在多个深度上利用接触相关线索。我们进一步将PSR实例化到视觉-语言-动作(VLA)模型中,形成PSR-VLA,并在六个真实世界的接触丰富操作任务上进行了评估。实验结果表明,PSR-VLA实现了91.7%的总体成功率,相比π0.5、ForceVLA-π0.5和ForceVLA2-π0.5分别提高了30.0、22.5和19.2个百分点。这些结果证明了所提出的PSR在力感知、接触丰富操作中的有效性。任务和稳定性测试的视频可在以下网址获取:此https URL。

英文摘要

Contact-rich manipulation requires policies to generate precise actions by reasoning over contact forces, robot configurations, and interaction histories beyond visual observations. Existing methods passively condition on force feedback rather than actively predicting future contact dynamics, limiting their ability to generate high-precision actions. To address this problem, we introduce Predictive Sensorimotor Representation (PSR) learning, a framework that learns a hierarchy of predictive representations from multimodal sensorimotor signals and integrates them into the action stream of a visuomotor policy. Specifically, during a pretraining stage, a multimodal Transformer is trained to learn a hierarchy of predictive representations by jointly forecasting future interaction dynamics. The learned hierarchy subsequently augments the action stream, enabling the resulting policy to exploit contact-relevant cues at multiple depths. We further instantiate PSR within a Vision-Language-Action (VLA) model, resulting in PSR-VLA, and evaluate it on six real-world contact-rich manipulation tasks. Experimental results show that PSR-VLA achieves 91.7% overall success, improving over $π_{0.5}$, ForceVLA-$π_{0.5}$, and ForceVLA2-$π_{0.5}$ by 30.0, 22.5, and 19.2 percentage points, respectively. These results demonstrate the effectiveness of the proposed PSR for force-aware, contact-rich manipulation. Videos of the tasks and stability tests are available at https://psr-vla.pages.dev/.

Comments7 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑