发表机构
University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ExoBridge通过人体肢体耦合,利用双臂协调行为,从裸手视频学习预测外骨骼运动与触觉状态,在1,215个演示中达到0.916的AUROC,建立了从人类视觉演示到机器人操作界面的桥接。
AI 中文摘要
人类视频为灵巧机器人学习提供了可扩展的经验来源,但在保持裸手交互的同时获取运动和触觉监督仍然具有挑战性。我们提出了ExoBridge,一个利用人体肢体耦合来学习从裸手视频到传感外骨骼的运动和触觉状态的桥接函数的框架。我们的核心思想是利用协调的双臂行为来连接未装备仪器的视觉来源与测量的操作界面。在数据采集过程中,一只手保持裸手并提供视觉观察,而另一只手佩戴外骨骼并提供同步的运动和触觉测量。这些成对的演示训练了一个时间视觉模型,仅从裸手视频预测指尖接触、连续触觉强度和相对编码器运动。外骨骼定义了一个中间状态空间,其运动坐标通过现有的校准与灵巧机器人手相关联。在四个操作任务的1,215个演示上的评估使用了留出采集会话,产生了合并的任何接触AUROC为0.916,触觉强度的Pearson相关系数为0.790。学习到的桥接函数还能从裸手视频预测外骨骼配置的相对变化。这些结果表明,人体肢体耦合可以将外骨骼测量转化为裸手视频的监督,建立了人类视觉演示与面向机器人的操作界面之间的学习桥接。
英文摘要
Human video offers a scalable source of experience for dexterous robot learning, but obtaining motion and tactile supervision while preserving bare hand interaction remains challenging. We present ExoBridge, a framework that leverages human limb coupling to learn a bridging function from bare hand video to the motion and tactile state of a sensorized exoskeleton. Our central idea is to use coordinated bimanual behavior to connect an uninstrumented visual source with a measured manipulation interface. During collection, one hand remains bare and provides visual observations, while the opposite hand wears the exoskeleton and supplies synchronized motion and tactile measurements. These paired demonstrations train a temporal visual model to predict fingertip contact, continuous tactile intensity, and relative encoder motion from bare hand video alone. The exoskeleton defines an intermediate state space whose motion coordinates are linked to a dexterous robot hand through existing calibration. Evaluation on 1,215 demonstrations across four manipulation tasks uses held out collection sessions and yields a pooled any contact AUROC of 0.916 and a Pearson correlation of 0.790 for tactile intensity. The learned bridge also predicts relative changes in exoskeleton configuration from bare hand video. These results demonstrate that human limb coupling can turn exoskeleton measurements into supervision for bare hand video, establishing a learned bridge between human visual demonstrations and a robot oriented manipulation interface.
Comments8 pages, 6 figures