Field Converter:面向足球转播的基于几何初始化的时间残差精化世界坐标球员姿态估计
Field Converter: Geometry-Initialized Temporal Residual Refinement for World-Grounded Player Pose Estimation from Soccer Broadcasts
- Institut de Biomécanique Humaine Georges Charpak (IBHGC), Arts et Métiers ParisTech(乔治·沙帕克人体生物力学研究所(IBHGC),巴黎高科国立高等工艺学校)
- French Football Federation (FFF)(法国足球联合会(FFF))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出Field Converter,利用球场几何初始化球员根部位置,并通过时间残差精化将世界坐标下姿态估计误差从49厘米降至10-11厘米,优于直接回归方法。
AI中文摘要:
从单目体育转播中恢复三维人体姿态仍然具有挑战性,尤其是当球员需要在共享的度量世界坐标系中进行定位,而不仅仅是相对于自身身体进行重建时。我们提出了Field Converter,一个基于几何初始化的时间残差框架,用于从标定的足球转播中估计世界坐标下的三维球员姿态。我们的方法首先利用相机和球场几何信息,通过射线-地面交点初始化球员根部位置,然后从姿态、图像、相机和几何线索中预测时间残差修正。在比赛不重叠的评估序列上,残差精化将根部误差从仅使用几何时的49厘米降低到使用逐帧MLP时的14厘米和使用TCN时的10厘米,而Transformer达到了相当的11厘米。最终的世界空间MPJPE达到13.2厘米,消融实验表明,残差预测明显优于直接全局根部回归,且时间上下文比具体的时间骨干网络更重要。失败分析进一步指出,空中运动是基于地面几何初始化的主要局限性。
英文摘要:
Recovering 3D human pose from monocular sports broadcasts remains challenging when players must be localized in a shared metric world coordinate system rather than only reconstructed relative to their own body. We introduce Field Converter, a geometry-initialized temporal residual framework for world-grounded 3D player pose estimation from calibrated soccer broadcasts. Our method first uses camera and pitch geometry to initialize the player root through ray-ground intersection, then predicts a temporal residual correction from pose, image, camera, and geometric cues. On match-disjoint evaluation sequences, residual refinement reduces root error from 49cm with geometry alone to 14cm with a frame-wise MLP and 10cm with a TCN, while a Transformer achieves a comparable 11cm. The resulting world-space MPJPE reaches 13.2cm, and ablations show that residual prediction clearly outperforms direct global-root regression while temporal context matters more than the specific temporal backbone. Failure analysis further identifies airborne motion as the main limitation of the ground-based geometric initialization.