发表机构
Seoul National University; RLWRLD(首尔大学; RLWRLD)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HOPE是一种基于单目视频的手-物体压力估计框架,通过将压力估计转化为手中心视频预测问题,结合顶点锚定视频Transformer等组件,在多基准测试中实现了良好泛化,可生成联合接触与压力预测。
AI 中文摘要
从视觉信息估计物理压力对于理解接触密集型手-物体交互至关重要。然而,现有的基于视觉的压力估计方法大多局限于平面表面和单张图像输入,难以应用于涉及不同物体的动态手-物体交互场景。我们转而将压力估计问题表述为以手为中心的视频预测问题,以单目视频作为输入。该表述直接在手部网格上预测随时间变化的逐顶点法向压力和接触情况,生成与物体形状和传感器布局无关的统一输出空间。基于此表述,我们提出了HOPE框架,该框架包含两个关键组件。首先,我们将触觉手套压力、平面传感器压力以及基于距离的手-物体接触标注提升至共享的手部顶点空间,使得在缺乏度量标签的情况下,裸手接触数据能够用于正则化压力学习。其次,我们引入了一种顶点锚定的视频Transformer,将每个顶点视为持久 token,随时间聚合视觉特征和手部姿态,并使用接触门控压力头确保在无接触时压力为零。在OpenTouch、PressureVisionDB以及手-物体接触基准上进行的实验,在物体压力、表面压力和接触监督的手-物体交互设置下对HOPE进行了验证。尽管主要使用来自戴手套手部视频的度量压力监督,HOPE仍能泛化到裸手的第一视角视频和野外视频,生成超出仅接触或平面压力基线范围的联合接触与压力预测。
英文摘要
Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estimation methods are largely limited to planar surfaces and single image input, making them difficult to apply to dynamic hand-object interaction with diverse objects. We instead formulate pressure estimation as a hand-centric video prediction problem with monocular video as input. This formulation predicts temporally evolving per-vertex normal pressure and contact directly on the hand mesh, yielding a unified output space independent of object shape and sensor layout. Building on this formulation, we propose \textbf{HOPE}, a framework with two key components. First, we lift tactile-glove pressure, planar-sensor pressure, and distance-based hand-object contact annotations into a shared hand vertex space, allowing bare-hand contact data to regularize pressure learning where metric labels are unavailable. Second, we introduce a vertex-anchored video transformer that treats each vertex as a persistent token, aggregates visual features and hand pose over time, and uses a contact-gated pressure head to enforce that pressure vanishes without contact. Experiments on OpenTouch, PressureVisionDB, and hand-object contact benchmarks validate HOPE across object-pressure, surface-pressure, and contact-supervised HOI settings. Despite using metric pressure supervision primarily from gloved-hand videos, HOPE generalizes to bare-hand egocentric and in-the-wild videos, producing joint contact and pressure predictions beyond the scope of contact-only or planar-pressure baselines.
Commentsproject page is at: https://subin6.github.io/page-hope