基于大型重建模型重构交互中的人体与物体
Reconstructing Humans and Objects in Interaction using Large Reconstruction Models
- University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出MILO框架,利用大型重建模型(LRMs)从单张图像重构三维人体-物体交互,在多个基准场景中精度优于现有方法。
AI中文摘要:
三维人体-物体交互(3D HOI)估计是三维计算机视觉的基础问题,应用于增强现实/虚拟现实(AR/VR)、机器人学和具身人工智能领域。然而,由于深度歧义、遮挡和物体形状的可变性,在三维空间中重构这些交互仍具有挑战性。现有方法主要关注重投影和接触约束,将参数化人体模型和物体模板拟合到二维图像中。在本文中,我们探索了一条不同的途径:提出了MILO框架,该框架利用大型重建模型(LRMs)的视觉能力,从单张图像中恢复详细的三维人体-物体交互。我们的关键观察是,LRMs提供了强大的几何支架,保留了人体与物体的相对排列和邻近线索,这显著简化了重构过程,将问题重新定义为解释LRMs生成的网格:将其分割为人体和物体组件,将参数化人体模型拟合到人体部分,若存在物体模板则可选地将其与物体部分对齐。MILO实现了较强的重构精度,在多个基准和交互场景中均优于现有基线方法。我们的代码可在该https URL获取。
英文摘要:
Estimation of Human-Object Interactions in 3D (3D HOI) is a fundamental problem in 3D computer vision with applications in AR/VR, robotics, and embodied AI. However, reconstructing these interactions in 3D remains challenging due to depth ambiguities, occlusions, and object shape variability. Existing approaches are primarily concerned with reprojection and contact constraints, fitting parametric human models and object templates to 2D images. In this paper, we explore a different avenue. We present MILO, a framework that leverages the visual capabilities of Large Reconstruction Models (LRMs) to recover detailed 3D human-object interactions from a single image. Our key observation is that LRMs provide a powerful geometric scaffold that preserves relative human-object arrangement and proximity cues. This significantly simplifies the reconstruction procedure, reframing the problem as interpreting the LRM mesh: we segment it into human and object components, fit a parametric body model to the human part, and optionally align an object template to the object part (if such a template is available). MILO achieves strong reconstruction accuracy and outperforms existing baselines across multiple benchmarks and interaction scenarios. Our code is available at https://ac5113.github.io/MILO.