ViTacPhys:基于人类视觉-触觉演示的感知物体物理属性的抓取方法
ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations
查看机构详情
- Xiaomi Robotics(小米机器人)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出ViTacPhys视觉-触觉框架,可从人类演示中估计物体物理属性,迁移至机器人后实现高抓取成功率,力曲线更贴合人类遥操作。
中文摘要 AI 辅助
近期,基于视觉的动作模型在复杂操作任务中展现出强大能力,但它们很少利用显式的物体物理属性来调整策略。本文提出ViTacPhys,这是一个视觉-触觉框架及数据采集系统,可从人类操作演示中估计物体质量、摩擦系数类别,以及连续的刚度值。该模型在60个刚性和可变形物体的数据上训练而成,结合了时序视觉-触觉建模、交叉注意力多模态融合,以及从视觉-语言模型导出的语义先验。在已见过的物体上,ViTacPhys实现了97.2%的质量分类准确率、98.8%的摩擦系数分类准确率,以及刚度平均绝对百分比误差(MAPE)为5.51%;在来自已知类别的未见过物体上,它实现了87.5%的质量准确率、97.5%的摩擦系数准确率,以及刚度MAPE为9.08%。我们通过有限的机器人遥操作数据、机器人风格的视频增强,以及动作匹配的人类演示,将ViTacPhys从人类领域迁移到机器人领域,并将其部署为自适应抓取的在线模块。由此得到的物理属性条件策略在分布内物体上实现了95.0%的总抓取成功率,在分布外物体上实现了83.4%的总抓取成功率。对于两种方法均成功抓取的分布外物体,该策略的力曲线比ACT生成的力曲线更符合人类遥操作的特征。这些结果证明,在实际自适应抓取中,显式估计并基于物体物理属性进行条件调整是可行的。
英文摘要
Recent vision-based action models have demonstrated strong capabilities in complex manipulation, but they rarely leverage explicit object physical properties to adapt their policies. We introduce ViTacPhys, a visual-tactile framework and data acquisition system that estimates object mass and friction-coefficient classes, together with continuous stiffness, from human manipulation demonstrations. Trained on data from 60 rigid and deformable objects, ViTacPhys combines temporal visual-tactile modeling, cross-attention multimodal fusion, and a semantic prior derived from a vision-language model. On seen objects, it achieves 97.2% mass classification accuracy, 98.8% friction-coefficient classification accuracy, and a stiffness mean absolute percentage error (MAPE) of 5.51%. On held-out objects from known categories, it achieves 87.5% mass accuracy, 97.5% friction-coefficient accuracy, and a stiffness MAPE of 9.08%. We transfer ViTacPhys from the human domain to the robot domain using limited robot teleoperation data, robot-style video augmentation, and human demonstrations with matched actions, and deploy it as an online module for adaptive grasping. The resulting physical-property-conditioned policy achieves total grasping success rates of 95.0% on in-distribution objects and 83.4% on out-of-distribution objects. For out-of-distribution objects successfully grasped by both methods, its force profiles are more consistent with human teleoperation than those produced by ACT. These results demonstrate the feasibility of explicitly estimating and conditioning on object physical properties for real-world adaptive grasping.