arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30026cs.CV

基于基础姿态模型的运动攀岩免训练抓点使用检测

Training-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models

Abu Bakar, Abdullah Aftab, Amir Hamza

首次发表
浏览论文内容

中文总结 AI 辅助

提出一种无需训练的攀岩抓点使用检测方法,利用基础姿态模型Sapiens的指尖和脚趾关键点,结合邻近测试和时序规则,在The Way Up数据集上达到90.2%的F1分数,优于现有方法。

中文摘要 AI 辅助

检测攀岩者使用哪些抓点以及何时使用,是运动攀岩自动化评分、动作分析和辅助系统的基础。现有方法训练任务专用模型,或改造2D姿态估计器,但其手部关键点位于手腕、脚部关键点位于脚踝,即与实际接触抓点的指尖和脚趾存在偏移,且手部在约一半的帧中被遮挡。我们证明,一个冻结的、现成的姿态基础模型就足够了:利用Sapiens的指尖和脚趾关键点,对标注抓点进行逐帧邻近测试、按肢体互斥以及短时时间持续性规则,我们无需任何攀岩专用训练即可检测抓点使用。在The Way Up数据集(22个视频,10名运动员,两条路线)上,我们的方法在留出分割上的事件F_1达到90.2%(留一参与者交叉验证下为89.8%),在所有22个视频上任意时间重叠下为79.9%,并在脚点上表现最佳(总体F_1为89.8%,留出为96.6%)。在相同协议下,它在每个时间阈值上都超过我们复现的YOLOv8-pose和ViTPose流程,且在严格时间条件下优势扩大。消融实验表明,两个直觉上有帮助的补充——密集基础特征变化门控和身体部位分割——都有害,说明最小化、仅关键点的设计才是此任务的正确选择。最后,从我们的自动预测计算出的标准教练统计与真实值高度吻合(攀爬时间Pearson r=1.00,节奏0.94),将普通单摄像头视频转化为无需仪器的可靠性能指标。

英文摘要

Detecting which holds a climber uses, and when, underpins automated scoring, movement analysis, and assistive systems for sport climbing. Existing approaches train task-specific models or repurpose 2D pose estimators whose hand keypoint sits at the wrist and foot keypoint at the ankle i.e. offset from the fingertips and toes that actually contact the holds, and whose hands are occluded in roughly half of all frames. We show that a frozen, off-the-shelf pose foundation model is sufficient: using the fingertip and toe keypoints of Sapiens, a per-frame proximity test against the annotated holds, per-limb mutual exclusion, and a short temporal-persistence rule, we detect hold usage without any climbing-specific training. On the The Way Up dataset (22 videos, 10 athletes, two routes), our method reaches an event F_1 of 90.2% on a held-out split (89.8% under leave-one-participant-out cross-validation) and 79.9% over all 22 videos at any temporal overlap, and performs best on footholds (F_1,89.8% overall, 96.6% held-out). Under an identical protocol it exceeds our reproductions of the YOLOv8-pose and ViTPose pipelines at every temporal threshold, with the margin widening under strict timing. An ablation shows that two intuitively helpful additions---dense foundation-feature change gating and body-part segmentation---both hurt, arguing that a minimal, keypoint-only design is the right one for this task. Finally, standard coaching statistics computed from our automatic predictions track ground truth closely (Pearson r=1.00 for climb time, 0.94 for pace), turning ordinary single-camera video into reliable performance metrics with no instrumentation.

发表机构

  • Virtual University of Pakistan(巴基斯坦虚拟大学)
  • Army Public School and Colleges(陆军公立学校与学院)
  • University of Trento(特伦托大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑