发表机构
University of Augsburg; National Institute of Informatics; University of Tokyo(奥格斯堡大学; 国立情报学研究所; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出EventNet两阶段流水线,利用2D关键点与变换器实现乒乓球视频中帧精确的球拍接触事件检测,在Latte-MV和TTHQ数据集上分别取得91.16%和73.08%的F1分数。
AI 中文摘要
本文解决了乒乓球视频中自动、帧精确事件检测的挑战。当前用于估计3D球轨迹和球旋转的方法通常要求关键事件(如球拍接触)已经预先被识别。这一要求使得这些方法难以应用于较长、未编辑的视频记录。为克服这一限制,我们提出了EventNet,一个两阶段流水线来检测关键事件:(1)提取双方球员的上半身姿态、球台角落和球心的2D关键点。一个小型关键点变换器将它们组合成紧凑表示,该表示对视角、光照和背景杂波的变化具有鲁棒性。(2)这些基于帧的表示的时间序列由变换器编码器处理,该编码器为每帧预测两个时间到事件值,指示当前帧距离下一次和上一次球拍接触的接近程度。一个新颖之处是新的时间余弦类目标信号。此外,我们通过3D重投影引入视角增强和帧率增强,以提高鲁棒性和泛化能力。我们广泛的消融研究深入揭示了各种架构和训练方面的重要性。实验结果表明,所提出的方法在Latte-MV数据集上实现了91.16%的F1分数和0.42的平均帧偏差(真实值与预测帧之间),在具有挑战性的TTHQ数据集上实现了73.08%的F1分数和1.16的平均帧偏差。总体而言,我们的工作表明,基于2D关键点的时间建模与我们的EventNet架构是乒乓球视频中自动事件检测的一种有前景且实用的方法。
英文摘要
This paper addresses the challenge of automatic, frame-accurate event detection in table tennis videos. Current methods for estimating 3d ball trajectories and ball spin typically require that key events, such as ball-racket contacts, have already been identified in advance. This requirement makes it difficult to apply these methods to longer, unedited video recordings. To overcome this limitation, we propose EventNet, a two-stage pipeline to detect key events: (1) 2d keypoints are extracted of the upper-body poses for both players, table corners and ball center. A small keypoint transformer combines them into a compact representation that is robust to changes in viewpoint, lighting, and background clutter. (2) The temporal sequences of these frame-based representations are processed by a transformer encoder that predicts two time-to-event values for each frame, indicating how close the current frame is to the next and previous ball-racket contact. One novelty is a new, temporal cosine-like target signal. Furthermore, we introduce viewpoint augmentation via 3D reprojection and frame-rate augmentation to improve robustness and generalization. Our extensive ablation study gives deeper insights into the importance of various architectural and training aspects. Experimental results show that the proposed approach achieves an F1 score of 91.16% and a mean frame deviation between ground truth and predicted frame of 0.42 on the Latte-MV dataset and 73.08% / 1.16 on the challenging TTHQ dataset. Overall, our work demonstrates that 2d keypoint-based temporal modeling with our EventNet architecture is a promising and practical approach for automatic event detection in table tennis videos.
CommentsAccepted at the 9th International ACM Workshop on Multimedia Content Analysis in Sports