发表机构
ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EgoPHI是从单目RGB图像与物体几何联合估计手和物体网格的密集3D接触图与力分布的方法,通过物理模拟扩充数据集,在多场景下实现了性能提升与泛化,推动了第一视角手-物体交互的物理推理。
AI 中文摘要
从第一视角视觉理解手-物体交互,对建模人如何与周围世界进行物理互动至关重要。然而,对基于物理的交互进行推理,除了定位接触外,还需要估计作用于手和物体上的力。我们提出EgoPHI,这是第一种从单目RGB图像和物体几何结构联合估计手和物体网格上的密集接触图与3D力分布的方法。为解决可扩展的力标注真值缺失问题,我们引入一种基于物理的模拟流程,该流程用密集的逐顶点力监督来扩充现有的手-物体数据集。随后,EgoPHI学习交互手与关节物体网格上的密集3D接触和力,将基于视觉的力估计扩展到图像空间或平面设置之外。我们在分布内和分布外基准上的评估显示,EgoPHI比现有方法提升了力估计性能,同时能泛化到未见过的数据集。为评估模拟到真实的迁移,我们构建了两个能捕捉密集物体接触和力大小的物理物体,并用它们记录了8名参与者在不同触摸和抓取类型下的交互数据集。我们的结果表明,EgoPHI在模拟、分布外和真实世界场景中都能恢复有意义的3D接触和力分布,推动第一视角手-物体理解从接触定位向基于物理的交互推理迈进。
英文摘要
Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet reasoning about physically grounded interaction requires estimating the forces acting on hands and objects, beyond localizing contact. We present EgoPHI, the first method that jointly estimates dense contact maps and 3D force distributions on hand and object meshes from a single monocular RGB image and object geometry. To address the lack of scalable ground-truth force annotations, we introduce a physics-based simulation pipeline that augments existing hand-object datasets with dense per-vertex force supervision. EgoPHI then learns dense 3D contact and force on interacting hand and articulated object meshes, extending vision-based force estimation beyond image-space or planar settings. Our evaluation on in-distribution and out-of-distribution benchmarks shows that EgoPHI improves force estimation over existing approaches while generalizing to unseen datasets. To evaluate sim-to-real transfer, we constructed two physical objects that capture dense object contact and force magnitude and used them to record a dataset of interactions from eight participants across diverse touch and grasp types. Our results demonstrate that EgoPHI recovers meaningful 3D contact and force distributions in simulated, out-of-distribution, and real-world settings, advancing egocentric hand-object understanding from contact localization toward physically grounded interaction reasoning.
CommentsAccepted by ECCV 2026