发表机构
MBZUAI; ETH Zürich; EPFL(穆罕默德·本·扎耶德人工智能大学; 苏黎世联邦理工学院; 洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PACT提出端到端联合学习框架,从单目视频同时估计人体姿态、接触与力,通过物理监督和ForceWall基准验证,优于分阶段方法。
AI 中文摘要
人体运动、环境接触和交互力受共同的物理规律支配,然而现有方法通常将视觉姿态重建与接触和力的估计分开处理。这种分离限制了联合推理,并可能在阶段之间传播误差。我们提出了PACT,一种端到端模型,可从单目视频中联合学习估计人体姿态、接触和接触力。我们的方法通过可学习的接触-力令牌和将视觉特征与世界空间运动相整合的时间变换器,增强了一个人体重建基础模型。联合预测头细化人体姿态并估计接触和力,而基于物理的监督则鼓励重建运动与交互力之间的一致性。为解决力标注稀缺的问题,我们开发了一种数据标注流程,该流程将接触标注与基于物理的运动和力优化相结合,从合成和真实世界视频中生成训练监督。我们还引入了一个真实世界的攀岩基准ForceWall,包含攀岩视频和从力传感器获得的相应地面真值接触力。实验表明,我们的方法在接触和力估计方面达到了最先进的性能,优于分阶段重建方法,并能泛化到训练分布之外的交互。这些结果支持端到端联合学习作为从视频恢复人体运动和物理交互的有效方法。
英文摘要
Human motion, environmental contacts, and interaction forces are governed by common physical laws, yet existing approaches typically separate visual pose reconstruction from contact and force estimation. This separation limits joint reasoning and can propagate errors between stages. We introduce PACT, an end-to-end model that jointly learns to estimate human pose, contacts and contact forces from monocular video. Our approach augments a human reconstruction foundation model with learnable contact-force tokens and a temporal transformer that integrates visual features with world-space motion. Joint prediction heads refine human poses and estimate contacts and forces, while physics-based supervision encourages consistency between the reconstructed motion and interaction forces. To address the scarcity of force annotations, we develop a data annotation pipeline that combines contact labeling with physics-based motion and force optimization, producing training supervision from synthetic and real-world videos. We also introduce a real-world climbing benchmark ForceWall with climbing videos and corresponding ground-truth contact forces obtained from the force sensors. Experiments demonstrate state-of-the-art contact and force estimation, outperforming staged reconstruction approaches and generalizing to interactions beyond the training distribution. These results support end-to-end joint learning as an effective approach to recovering human motion and physical interactions from video.
CommentsProject page: https://rihat99.github.io/PACT/