arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17097cs.CV

HarmoHOI:用于多视图手-物体交互合成的外观与3D运动协调

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

Lingwei Dang, Juntong Li, Zonghan Li, Hongwen Zhang, Liang An, Wei Min, Yebin Liu, Qingyao Wu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对多视图手-物体交互合成难题,提出HarmoHOI统一扩散框架,用多视图扩散Transformer混合模型联合建模视频与点轨迹,结合全局运动对齐扩散确保一致性,采用混合数据学习策略,实现多视图HOI高质量合成。

中文摘要 AI 辅助

手-物体交互(HOI)合成是动画制作和具身人工智能的基石。尽管视频基础模型有强大先验,但由于手部动作复杂和遮挡,多视图一致的HOI合成仍具挑战性。我们提出HarmoHOI,一个统一的扩散框架,可联合并协调地生成同步多视图HOI视频和全局对齐的3D点轨迹。核心见解是强大的多视图一致性根本上需要全局对齐的3D几何和运动。为此,我们提出多视图扩散Transformer混合模型来共同建模RGB视频和3D点轨迹。通过将点轨迹表示为伪视频,使3D几何信号与基础模型的2D潜在空间对齐,最小化域差距并简化先验适应。为进一步确保几何一致性,引入全局运动对齐扩散,将粗糙点轨迹细化为度量尺度、全局对齐的3D轨迹。HarmoHOI能在去噪过程中实现2D外观和3D运动的实时协同进化。为克服多视图HOI数据稀缺问题,采用混合数据课程学习策略,成功将单视图数据的通用先验转移到同步多视图生成中。实验结果表明HarmoHOI在视觉质量、运动合理性和多视图几何一致性方面达到了当前最优性能。

英文摘要

Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to complex hand motions and occlusions. We present HarmoHOI, a unified diffusion framework that jointly and harmoniously generates synchronized multi-view HOI videos and globally aligned 3D point tracks. Our core insight is that robust multi-view consistency fundamentally requires globally aligned 3D geometry and motion. To this end, we propose a Mixture of Multi-view Diffusion Transformer that co-models RGB videos and 3D point tracks. By representing point tracks as pseudo-videos, we align 3D geometric signals with the 2D latent space of foundation models, thereby minimizing the domain gap and easing adaptation of priors. To further ensure geometry consistency, we introduce Global Motion Aligning Diffusion, which refines coarse point tracks into metric-scale, globally aligned 3D trajectories. HarmoHOI enables on-the-fly co-evolution of 2D appearance and 3D motion during denoising. To overcome the scarcity of multi-view HOI data, we employ a hybrid data curriculum learning strategy that successfully transfers generic priors from single-view data to synchronized multi-view generation. Experimental results show that HarmoHOI achieves state-of-the-art performance in visual quality, motion plausibility, and multi-view geometric consistency. Project page available at https://droliven.github.io/HarmoHOI_project.

发表机构

  • South China University of Technology(华南理工大学)
  • Beijing Normal University(北京师范大学)
  • Tsinghua University(清华大学)
  • Shadow AI(影谱科技)

机构由 AI 辅助整理,请以论文原文为准。

↑