VisForce:面向目标条件灵巧操作中当前力与期望力的视觉锚定
VisForce: Visual Grounding of Current and Desired Forces for Goal-Conditioned Dexterous Manipulation
浏览论文内容
中文总结 AI 辅助
针对灵巧手操作中力与视觉位置对应困难的问题,提出VisForce方法,将当前与期望力在指尖位置视觉锚定,结合目标条件交叉注意力生成力感知动作,实验验证了其有效性。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型已成为通用机器人操作策略。然而,在灵巧手操作中,接触力通常作为独立状态或特定于力的表示提供,这使得难以明确表示力与其对应视觉位置之间的空间对应关系。在这项工作中,我们提出了VisForce,它在其对应的指尖位置对当前力和期望力进行视觉锚定。VisForce在当前腕部图像和任务特定的目标图像上渲染当前和期望的视觉力提示,并通过目标条件交叉注意力结合这两种表示以生成力感知动作。我们使用配备RH56F1灵巧手的真实UR10机器人,通过力条件抓取和三个多阶段操作任务评估VisForce。在力条件抓取实验中,VisForce随着期望力的增加表现出一致的抓握力响应,并分别对鸡蛋和牙膏管实现了70%和80%的抓取-提升成功率。它还在杯子插入/倒瓶、钳子辅助面包转移和滑移调节的轴孔插入任务中分别达到了70%、55%和40%的最终成功率。这些结果表明,指尖对齐的视觉力表示可有效用于基于VLA的灵巧手操作中的力感知条件设定。
英文摘要
Vision-Language-Action (VLA) models have emerged as general-purpose robotic manipulation policies. However, in dexterous hand manipulation, contact forces are typically provided as separate states or force-specific representations, making it difficult to explicitly represent the spatial correspondence between force and their corresponding visual locations. In this work, we propose VisForce, which visually grounds the current and desired forces at their corresponding fingertip locations. VisForce renders current and desired visual force cues on the current wrist image and a task-specific goal image, and combines the two representations through goal-conditioned cross-attention to generate force-aware actions. We evaluate VisForce using a real UR10 robot equipped with an RH56F1 dexterous hand through force-conditioned grasping and three multi-stage manipulation tasks. In force-conditioned grasping experiments, VisForce exhibited a consistent grip-force response as the desired force increased, and achieved grasp-and-lift success rates of 70% and 80% for an egg and a toothpaste tube, respectively. It further achieved final success rates of 70%, 55%, and 40% on cup insertion/bottle pouring, tong-assisted bread transfer, and slip-modulated peg-in-hole, respectively. These results show that fingertip-aligned visual force representations can be effectively used for force-aware conditioning in VLA-based dexterous hand manipulation.