发表机构
Institute of Science Tokyo; Carnegie Mellon University(东京科学研究所; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
因缺乏高质量数据集,人宠交互研究不足。本文提出InterPet4D多模态数据集及InterPetMoGen框架,通过同步多视图系统记录交互并标注,含多类型数据,所提模型FID得分11.21,优于基线,有效模拟人宠交互。
AI 中文摘要
由于缺乏高质量大规模数据集,人宠交互估计和生成仍未得到充分探索。我们展示了InterPet4D,这是首个捕捉人类与狗自然交互的多模态数据集。利用同步多视图捕捉系统,记录人犬服从任务并为人和狗提供注释,包括多视图和自我中心视频、分割、2D和3D关键点、网格及音频轨道。该数据集由680万帧组成,来自11个品种的13只狗与23名人类参与者的交互。我们还引入了InterPetMoGen框架用于人宠交互运动生成。我们提出的模型FID得分为11.21,显著优于Seq2Seq和DiT基线,证明了InterPet4D对模拟现实人宠交互的有效性。
英文摘要
Human-pet interaction estimation and generation remain underexplored due to the absence of a high-quality large-scale dataset. We present InterPet4D, the first multimodal dataset capturing natural interactions between humans and dogs. Using a synchronized multi-view capture system, we record human-dog obedience tasks and provide annotations for both humans and dogs, including multi-view and egocentric videos, segmentations, 2D and 3D keypoints, meshes, and audio tracks. InterPet4D consists of 6.8 million frames collected from 13 dogs of 11 breeds interacting with 23 human participants. We further introduce the InterPetMoGen framework for human-pet interaction motion generation. Our proposed model achieves an FID score of 11.21 and substantially outperforms the Seq2Seq and DiT baselines, demonstrating the effectiveness of InterPet4D for modeling realistic human-pet interactions.