发表机构
University of Science, VNU-HCM; Vietnam National University, Ho Chi Minh City(胡志明市国家大学科学大学; 胡志明市越南国家大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对单帧转播画面中运动员定位的极端尺度差异难题,提出自上而下框架,含边界感知自适应分块与RTMPose-X适配的两点估计,在公开测试集LocSim达97.44、mAP达0.9128,性能超基准21%以上。
AI 中文摘要
由于超高分辨率图像中存在极端尺度差异,从单帧转播画面中对运动员进行精确的世界坐标定位本质上具有挑战性。本文提出一种用于从单张已校准帧实现度量尺度运动员定位的自上而下框架,该方法包含三项核心贡献:第一,提出边界感知自适应分块(Boundary-Aware Adaptive Tiling),这是对标准切片推理的语义引导扩展,通过基于粗略边界框预测迭代扩展分块边界,系统确保完整包含目标对象,通过轻量级流水线调整且不修改架构,有效缓解边界分割伪影;该方法大幅降低了极端尺度差异下的召回率下降,使我们能将透视畸变确定为剩余定位误差的主要来源。第二,将RTMPose-X架构适配为专门的两点(骨盆与地面投影)估计器,采用针对该几何耦合点对优化的重新设计门控注意力单元(Gated Attention Unit),随后通过相机校准的光线投射将2D地面投影确定性提升至世界坐标。在公开测试集上,该方法达到97.44的LocSim分数和0.9128的mAP,较基准方法提升超过21%,为高分辨率尺度差异问题建立了鲁棒解决方案。
英文摘要
Accurate world-coordinate localization of athletes from single-frame broadcast footage is inherently challenging due to extreme scale disparities in ultra-high-resolution imagery. In this paper, we propose a top-down framework for metric-scale athlete localization from a single calibrated frame. Our approach centers on three key contributions. First, we propose Boundary-Aware Adaptive Tiling, a semantics-guided extension of standard sliced inference. By iteratively expanding tile boundaries based on coarse bounding-box predictions, it systematically ensures full object containment, effectively mitigating boundary-splitting artifacts through a lightweight pipeline adaptation without architectural modifications. By substantially mitigating recall degradation under extreme scale variance, Boundary-Aware Adaptive Tiling enables us to isolate perspective distortion as the primary source of residual localization error. Second, we adapt the RTMPose-X architecture into a specialized two-keypoint estimator (pelvis and ground projection), employing a reformulated Gated Attention Unit optimized for this geometrically coupled point pair, and then deterministically lift the 2D ground projections into world coordinates via camera-calibrated ray casting. On the public test set, our method achieves a LocSim score of 97.44 and an mAP of 0.9128, outperforming the baseline by over 21 \% and establishing a robust solution for high-resolution scale variance.