PL-NBA:支持多种视觉理解任务的控球级通用篮球视频数据集
PL-NBA: A Possession-level Universal Basketball Video Dataset Supporting Multiple Visual Understanding Tasks
浏览论文内容
中文总结 AI 辅助
本文构建首个控球级篮球视频数据集PL-NBA,包含11000个进攻控球片段及31567个注释事件,在四项视觉理解任务上验证其挑战性,为体育视频理解提供基准。
中文摘要 AI 辅助
近年来,体育视觉理解已成为计算机视觉领域的热门话题。现有大多数篮球视频数据集以单一动作或活动作为样本,既无法保留比赛事件的时间连续性,也无法支持动作预测等复杂任务。为解决该问题,本文构建了首个控球级篮球视频数据集(PL-NBA),每个样本由一段完整的NBA进攻控球回合组成。PL-NBA从60场NBA比赛中收集,包含11000个有效进攻控球片段和31567个带注释的事件,注释内容涵盖球员姓名、描述文本、事件类型和时间戳。每个视频片段包含多个事件并保留事件的连续性,有助于战术分析。本文在多个视觉理解任务上开展实验,包括事件识别、视频字幕生成、时间动作定位和动作预测。实验结果表明,现有方法在上述四项任务上的性能有限,证明PL-NBA是体育视频理解领域的一个具有挑战性的基准。
英文摘要
Visual understanding in sports has emerged as a hot topic in computer vision in recent years. Most existing basketball video datasets adopt single action or activity as sample, which can neither preserve the temporal continuity of game events nor support complex tasks such as action anticipation. To address this issue, this paper constructs the first possession-level basketball video dataset (PL-NBA), in which each sample is composed of a complete NBA offensive possession. Collected from 60 NBA games, PL-NBA contains 11,000 valid offensive possession clips and 31,567 annotated events with player names, captions, event types and timestamps. Each video clip includes multiple events and preserves the continuity of events, which is helpful for analysis of tactic. Experiment is conducted on multiple visual understanding tasks, including event recognition, video captioning, temporal action localization and action anticipation. Experimental results show that existing methods achieve limited performance on above four tasks, demonstrating that PL-NBA is a challenging benchmark for sports video understanding.
发表机构
- Beijing University of Technology(北京工业大学)
- Nanjing University of Science and Technology(南京理工大学)
- Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。