发表机构
Peking University; The Hong Kong University of Science and Technology; Tencent; East China Normal University; Sun Yat-sen University; Zhejiang University(北京大学; 香港科技大学; 腾讯; 华东师范大学; 中山大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlowHMR提出将视频运动捕捉视为视频条件下的运动生成问题,利用流匹配模型和GRPO优化,结合保真度与跟踪奖励,在Wild-4K数据集上实现了82.47%的物理跟踪成功率,优于现有方法。
AI 中文摘要
我们提出了FlowHMR,一个从单目视频中恢复物理上合理的全局3D人体运动的框架。以往基于学习的方法通常直接从视频回归人体运动,并用几何监督训练网络。然而,从单目视频恢复人体运动在深度上本质上是模糊的,直接回归往往趋向于平均化的解决方案。此外,恢复的运动不能保证在物理上是合理的,因此基于物理的跟踪常常失败。为了解决这些挑战,我们将视频运动捕捉表述为一个视频条件下的运动生成问题,并首先为此任务预训练一个流匹配模型。给定输入视频,预训练模型生成多样的运动候选,但并非所有候选都忠实于视频或可物理跟踪。因此,我们使用组相对策略优化(GRPO)并配以两个奖励对模型进行后训练。一个保真度奖励鼓励与输入视频的一致性。一个跟踪奖励偏好那些物理控制器能够成功跟踪的运动。这些奖励共同改变了模型的输出偏好,使得后训练模型在保持对输入视频忠实的同时,产生更物理上合理的运动。我们进一步引入了Wild-4K,一个包含约4000个互联网视频的大规模多样化数据集,用于评估野外人体运动恢复。在Wild-4K上的定性和定量实验表明,我们的方法在整体运动保真度上优于最先进的方法,并实现了82.47%的物理跟踪成功率,而最强基线GVHMR为62.82%。
英文摘要
We present FlowHMR, a framework for recovering physically plausible global 3D human motion from monocular video. Previous learning-based methods typically regress human motion directly from video and train the network with geometric supervision. However, recovering human motion from monocular video is inherently ambiguous in depth, and direct regression tends to collapse toward an averaged solution. Moreover, the recovered motions are not guaranteed to be physically plausible, so physics-based tracking of them often fails. To address these challenges, we formulate video motion capture as a video-conditioned motion generation problem and first pretrain a flow matching model for this task. Given an input video, the pretrained model generates diverse motion candidates, but not all of them are faithful to the video or physically trackable. We therefore post-train the model using Group Relative Policy Optimization (GRPO) with two rewards. A fidelity reward encourages consistency with the input video. A tracking reward favors motions that a physics-based controller can track successfully. Together, these rewards shift the model's output preference, so the post-trained model stays faithful to the input video while producing more physically plausible motion. We further introduce Wild-4K, a large and diverse dataset of about 4K internet videos, for evaluating human motion recovery in the wild. Qualitative and quantitative experiments on Wild-4K show that our method outperforms state-of-the-art methods in overall motion fidelity and achieves a physical tracking success rate of 82.47%, compared with 62.82% for the strongest baseline, GVHMR.
CommentsProject page: https://flowhmr.github.io/ Code: https://github.com/flowhmr/flowhmr