发表机构
DGIST; Seoul National University(大邱庆北科学技术院; 首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VICON提出一种视觉-惯性-接触融合的手物跟踪框架,在严重遮挡下同步捕捉手物运动与接触信息,显著降低失败帧率并构建公开数据集。
AI 中文摘要
学习灵巧操作受益于人类演示数据集,这些数据集捕捉多样且自然的手物交互。特别是,接触点和接触力提供了关于交互位置和交互强度的监督信息,这些信息无法仅通过运动轨迹完全捕捉。然而,同时捕捉手和物体运动、接触点及接触力的方法仍然有限。此外,手物交互导致的严重遮挡对手和物体的准确跟踪构成挑战。为解决这些限制,我们提出了一个基于视觉-惯性-接触的手物跟踪(VICON)框架。该框架在操作过程中整体捕捉手和物体的运动以及接触信息,即使在严重遮挡下也能实现。首先,我们采用视觉-惯性手套和RGB-D相机进行准确的手部跟踪,并重新设计手套以集成接触传感。具体而言,基于人类抓取频率将力敏电阻(FSRs)放置在手套上,以同步记录接触状态和校准的法向力。其次,无需预先存在的CAD模型,我们使用RGB-D图像和从单目视频重建的网格来估计物体姿态。我们提出了基于因子图的物体轨迹估计方法,该方法融合了在手物遮挡下按可见性加权的物体姿态估计、FSR测量和手部运动先验。在涉及五个物体的40次运动捕捉会话中,VICON实现了2.5%的失败帧率,而基线方法的失败帧率为50.9%-64.6%,在遮挡下的中位误差为3.9毫米和3.0度。使用VICON,我们构建了一个包含同步手物运动、接触点和法向力的数据集,并将在该https URL上公开发布覆盖10个物体类别的扩展数据集。
英文摘要
Learning dexterous manipulation benefits from human demonstration datasets that capture diverse and natural hand-object interactions. In particular, contact points and forces provide supervision on where and how strongly to interact, which cannot be fully captured by motion trajectories alone. However, methods for jointly capturing hand and object motion, contact points, and forces remain limited. Moreover, severe occlusion from hand-object interaction challenges accurate tracking of both hands and objects. To address these limitations, we present a Visual-Inertial-CONtact based hand-object tracking (VICON) framework. It holistically captures both hand and object motion along with contact information during manipulation, even under severe occlusion. First, we adopt a visual-inertial glove and an RGB-D camera for accurate hand tracking, and redesign the glove to incorporate contact sensing. Specifically, force-sensitive resistors (FSRs) are placed on the glove based on human grasp frequency to synchronously record contact states and calibrated normal forces. Second, without requiring pre-existing CAD models, we estimate object poses using RGB-D images and a mesh reconstructed from a monocular video. We propose factor-graph-based object trajectory estimation that fuses object-pose estimates weighted by visibility under hand-object occlusion, FSR measurements, and a hand-motion prior. Across 40 motion-capture sessions with five objects, VICON achieves a 2.5% failed-frame rate compared with 50.9-64.6% for the baselines, with median errors of 3.9 mm and 3.0 degrees under occlusion. Using VICON, we construct a dataset containing synchronized hand-object motion, contact points, and normal forces, and will publicly release an expanded dataset covering 10 object categories at https://github.com/VICON-dataset/dataset.
Comments9 pages, 6 figures, 4 tables