GraspHOI:从单张野外图像重建带手指级抓取的全身3D人-物体交互
GraspHOI: Full-Body 3D Human-Object Reconstruction with Finger-Level Grasps from a Single In-the-Wild Image
- Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
GraspHOI是首个从单张野外图像重建带手指级抓取的全身3D人-物体交互的框架,通过显式优化手指关节实现合理接触,在多基准测试中性能优于现有方法,代码将公开。
AI中文摘要:
现有的单目全身3D人-物体交互(HOI)方法未将显式的手指级抓取优化与类别无关的物体重建相结合。尽管这些方法能生成看似合理的人体-物体配置,但它们的手指可能会漂浮在物体表面或穿透物体,而非形成抓取姿态。本文提出GraspHOI,这是首个从单张图像重建全身3D HOI,同时显式优化手指关节以匹配重建物体的框架。GraspHOI直接恢复物体几何结构,无需预定义网格或固定类别词汇表,它分别重建人体、手部和物体,通过基于深度的配准和图像空间对齐将它们在度量相机空间中对齐;基于遮挡的手掌对应关系使物体贴合抓取的手部,接触感知优化则细化手臂和手指关节,以形成表面接触且无过度穿透。在四个基准测试和六个基线方法上,GraspHOI在人体-物体相对放置、手部精度和接触合理性方面均有所提升,完整的代码将被发布。
英文摘要:
Existing monocular full-body 3D human-object interaction (HOI) methods do not combine explicit finger-level grasp optimization with category-agnostic object reconstruction. Despite plausible body-object configurations, their fingers may float from or penetrate objects instead of forming a grasp. We present GraspHOI, the first framework that reconstructs a full-body 3D HOI from a single image while explicitly optimizing finger articulation against the reconstructed object. GraspHOI recovers object geometry directly, without predefined meshes or a fixed category vocabulary. It reconstructs the body, hands, and object separately, aligning them in metric camera space via depth-based registration and image-space alignment. Occlusion-aware palmar correspondences seat the object against the grasping hand, and contact-aware optimization refines arm and finger articulation to form surface contact without excessive penetration. Across four benchmarks and six baselines, GraspHOI improves relative human-object placement, hand accuracy, and contact plausibility. Full pipeline code will be released.