AI 中文总结
PhotoHOI从单张RGB照片与开放词汇指令合成3D手-物体交互,通过视觉-语言模型解析任务规格、规划碰撞感知轨迹、学习接触与抓取先验,在GRAB、H2O数据集及真实照片上表现优异。
AI 中文摘要
手-物体交互(HOI)是一种基础人类行为,在增强现实/虚拟现实(AR/VR)、数字人、具身交互等领域有广泛应用。现有方法通常需要预设的物体几何、物体轨迹或特定任务条件,限制了其在自然真实世界输入中的应用。为解决这一问题,本文研究了一个更具实用性的问题:从单张RGB照片和开放词汇语言指令合成3D手-物体交互序列,并提出了PhotoHOI方法。PhotoHOI首先利用视觉-语言模型将输入图像和指令解析为结构化任务规格,包含交互物体、目标区域和空间关系;随后基于恢复的物体状态、支撑关系及周围场景几何,恢复紧凑的任务相关3D场景并规划平滑的碰撞感知物体轨迹。为合成可泛化到真实世界照片和未见物体的手部运动,PhotoHOI从大规模 affordance 和 HOI 数据中学习可迁移的任务条件接触及接触条件抓取先验,抓取还会在学习到的潜在空间中进一步优化,将优化过程约束在合理的手部姿态流形内。在GRAB和H2O数据集上的实验表明,与代表性基线相比,PhotoHOI的接触质量有所提升,穿透现象减少;在真实世界照片上的结果进一步显示其具有更高的任务成功率和场景一致性,且能泛化到未见物体和开放词汇指令。
英文摘要
Hand-object interaction (HOI) is a fundamental human behavior with broad applications in AR/VR, digital humans, and embodied interaction. Existing methods typically require predefined object geometry, object trajectories, or task-specific conditions, limiting their use with natural real-world inputs. To address this, we study a more practical problem of synthesizing 3D hand-object interaction sequences from a single RGB photograph and an open-vocabulary language instruction, and introduce PhotoHOI. PhotoHOI first uses a vision-language model to parse the input image and instruction into a structured task specification, including the interaction object, target region, and spatial relation. It then recovers a compact task-relevant 3D scene and plans a smooth collision-aware object trajectory based on the recovered object states, support relations, and surrounding scene geometry. To synthesize hand motion that generalizes to real-world photographs and unseen objects, it learns transferable task-conditioned contact and contact-conditioned grasp priors from large-scale affordance and HOI data. The grasp is further refined in a learned latent space, constraining the optimization to a plausible hand-pose manifold. Experiments on GRAB and H2O demonstrate improved contact quality and reduced penetration over representative baselines. Results on real-world photographs further demonstrate higher task success and scene consistency, together with generalization to unseen objects and open-vocabulary instructions.