发表机构
Skolkovo Institute of Science and Technology(斯科尔科沃科学技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ZeroTouch提出触觉监督的视觉接触估计框架,利用腕部RGB预测接触变形和力,无需部署触觉硬件,显著提升抓取成功率。
AI 中文摘要
可靠的机器人抓取受益于对不断演变的物理交互的估计以及选择依赖于抓取的压缩目标。触觉传感器提供直接的交互测量,但需要在部署时配备专用硬件。我们提出了ZeroTouch,一个触觉监督框架,它从腕部RGB观测、夹爪状态和局部重力方向预测密集接触变形、瞬时六轴力/力矩以及依赖于抓取的期望压缩目标。触觉测量仅在训练期间用作特权监督,部署时不需要。在完整验证集上,完整架构将法向力平均绝对误差从仅状态基线的2.017 N降低到0.531 N。在每种条件20次试验的物理评估中,ZeroTouch在未见物体上达到95%的成功率,在已见物体/未见抓取条件下达到80%,在内容/负载变化下达到90%。在相同的评估协议下,OpenVLA分别达到25%、40%和55%,而SmolVLA分别达到10%、25%和35%。
英文摘要
Reliable robotic grasping benefits from estimating the evolving physical interaction and selecting a grasp-dependent compression target. Tactile sensors provide direct interaction measurements but require dedicated hardware at deployment. We introduce ZeroTouch, a tactile-supervised framework that predicts dense contact deformation, the instantaneous six-axis wrench, and a grasp-dependent desired compression target from wrist RGB observations, gripper state, and local gravity direction. Tactile measurements are used only as privileged supervision during training and are not required at deployment. On the full validation set, the complete architecture reduces normal-force MAE from 2.017 N for a state-only baseline to 0.531 N. In physical evaluation with 20 trials per condition, ZeroTouch achieves 95% success on an unseen object, 80% in a seen-object/unseen-grasp condition, and 90% under a content/load shift. Under the same evaluation protocol, OpenVLA achieves 25%, 40%, and 55%, while SmolVLA achieves 10%, 25%, and 35%, respectively.