发表机构
School of AI, Shanghai Jiao Tong University; Rutgers University; The University of Hong Kong; QuicRobot(上海交通大学人工智能学院; 罗格斯大学; 香港大学; QuicRobot)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EgoFound3R提出统一端到端模型,在世界空间以度量尺度重建手部几何并预测逐点交互属性,通过三种设计降低推理成本,在多个基准上显著降低MPJPE并提升吞吐量。
AI 中文摘要
自我中心视频已成为具身模型监督的主要来源,其价值在于恢复世界坐标中的手部运动,而相机运动和手部遮挡使得这一任务变得困难。现有重建流程通常将手部和场景估计分开,将交互属性留给单独的任务特定模型,并且每个视频调用多个模型,因此没有先前的重建模型估计这些属性,吞吐量成为大规模标注的实际约束。因此,我们引入EgoFound3R,一个统一的端到端模型,以与场景共享的度量尺度估计世界空间中的手部几何,并预测逐点交互属性,包括可见性、接触和距离。该模型整合了三种设计:(i)结构化手部提示,将预训练的几何先验转移到世界空间手部重建;(ii)显式手部表示,解码手部几何和交互属性;(iii)共享参数的多速率设计,降低推理成本。这些设计共同在一次前向传递中预测手部几何和逐点属性。在OakInk-v2、TACO和HOI4D上,EgoFound3R将平均每关节位置误差(MPJPE)相较于先前方法分别降低了43.2%、22.4%和11.6%,并在同一传递中预测逐点接触和距离以及几何,同时实现了约6倍的更高吞吐量。
英文摘要
Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to separate task-specific models, and invoke several models per video, so no prior reconstruction model estimates these attributes and throughput becomes a practical constraint on large-scale annotation. We therefore introduce EgoFound3R, a unified end-to-end model that estimates world-space hand geometry in a metric scale shared with the scene, and predicts point-wise interaction attributes, including visibility, contact, and distance. The model integrates three designs: (i) structured hand prompts that transfer pretrained geometric priors to world-space hand reconstruction; (ii) an explicit hand representation that decodes hand geometry and interaction attributes; and (iii) a shared-parameter multi-rate design that lowers inference cost. Together, these designs predict hand geometry and point-wise attributes in one pass. On OakInk-v2, TACO, and HOI4D, EgoFound3R reduces the mean per-joint position error (MPJPE) by 43.2%, 22.4%, and 11.6% over previous methods and predicts point-wise contact and distance alongside the geometry in the same pass, while attaining approximately 6x higher throughput.