arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RLHND:视频基础模型作为物理接地的手部追踪器用于机器人学习

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

Seungjun Moon, Subin Jeon, Sangwoo Kim, Hanbyul Joo, Jinwoo Shin

arXiv 2610.09455首次发表:更新:

发表机构

RLWRLD; KAIST; Seoul National University(RLWRLD; 韩国科学技术院; 首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RLHND,利用视频基础模型从第一人称视频联合估计手部姿态与触觉信息,实现物理一致的手部追踪,提升机器人学习效果。

AI 中文摘要

近年来,利用人类视频数据集进行机器人策略训练的方法日益普遍。然而,大多数现有手部追踪器从裁剪帧中回归姿态,对手部运动和物体交互的先验知识有限,导致估计结果不准确且在物理上不一致。此外,缺乏物理线索(如接触和力)限制了人类视频在机器人策略训练中的使用。为此,我们提出RLHND,一种基于视频基础模型的手部追踪模型,能够从单目第一人称视频中联合估计手部姿态和真实的触觉信息。RLHND通过干净潜变量条件化,将预训练的Cosmos 3视频扩散骨干网络转变为确定性的片段级特征提取器,将其在手部运动和手-物体交互方面的学习先验带入追踪任务。对于姿态估计,RLHND(i)预测具有解剖学上合理关节角度的姿态,并且(ii)支持对形状参数的可选条件化,以在同一个视频内甚至由同一演员录制的不同视频中保持手部形状一致。对于触觉估计,一个独立的触觉专家流,在姿态流冻结的情况下训练,预测手部表面的密集接触和力。我们进一步采用基于LBS的特征传播,以实现无需昂贵的逐顶点注意力的逐顶点特征提取。RLHND在多个基准数据集上实现了姿态估计的最先进性能,同时在接触和力估计方面也达到了最先进水平。此外,我们通过重定向结果和真实世界机器人实验展示了RLHND在机器人学习中的实用性。代码将在该https URL公开提供。

英文摘要

Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at https://seungjun-moon.github.io/rlhnd/.

Comments34 pages, 15 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑