AI 中文总结
本综述系统梳理大基础模型在手物交互(HOI)领域的应用,建立基础模型子先验分类,分析其在HOI任务与机器人学习中的作用,总结数据集与评估协议并展望未来方向。
AI 中文摘要
手物交互(HOI)建模仍具挑战性,因为它需要在严重视觉不确定性下对手部关节运动、物体几何、接触、语义和动力学进行联合推理。基础模型引入了从大规模跨域数据中学习到的可迁移先验知识,为解决这些挑战提供了超越特定任务数据和模型的新途径。然而,快速增长的相关文献仍处于碎片化状态,现有研究通常仅将这些方法简单描述为“使用大模型”,未系统刻画引入了何种知识、其在HOI流程中的何处介入,以及它有助于减少哪些HOI不确定性。本综述首次对HOI相关的基础模型先验进行系统回顾,将文献组织为涵盖重建与生成的6项HOI任务;更重要的是,建立了8种基础模型子先验的分类体系,分为几何、语义和视觉三类:几何先验包括形状检索、形状重建和空间重建;语义先验包括语义 grounding 和语言推理;视觉先验涵盖视觉表示、图像生成和视频生成。基于该分类,系统分析了不同先验在HOI流程和任务中的表示、注入与适配方式。除了基础模型赋能HOI的方式,还进一步研究了HOI衍生知识在机器人学习中的应用,包括人类数据预训练、人到机器人的技能迁移以及HOI到机器人的数据生成。最后,总结了数据集和评估协议,讨论了更具泛化性的HOI系统的局限性和未来方向,为支持长期进展,整理了一个持续聚合新兴方法和基准的在线仓库。
英文摘要
Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models'' without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.