发表机构
Sports Vision, Inc.(Sports Vision, Inc.)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对业余单摄像头体育视频,本研究以排球为例,比较提示、探测、训练与标注四种方法,发现无单一范式全面胜出,身份识别为共同难点,并探讨各方法的成本效益与迁移价值。
AI 中文摘要
视频理解通常在精心策划、单一演员或专业拍摄的片段上进行基准测试,而在此类测试中取得的高分常被视为模型足够稳健、可部署于实际应用的证据。业余团队运动是一个有用且基本未经检验的场所,用以检验这一假设:仅在美国,2024-25学年就有超过八百万学生参加学校体育运动,其中几乎所有的比赛都仅由一台固定摄像头拍摄,画面中多名候选运动员拥挤在一起,没有操作员或第二视角可供依赖。以排球为测试案例,我们探究在如此混乱的镜头下,通用视频和世界模型基准上的强劲性能是否能转化为可靠的、逐球员的归因,通过从寻找比赛边界到命名谁做了什么等一系列任务,将视频转化为统计数据。我们在每个阶段评估四种方法(对前沿视觉-语言模型进行提示和智能体推理、使用小型训练专家的经典计算机视觉、自监督视频世界模型以及人工标注),共涉及66场业余比赛,包含46,648个由人工标注的接触事件,拍摄条件为现有公开基准所未曾使用的。没有单一范式在每个阶段都胜出,而静态、单帧的计算机视觉在任何涉及运动或身份的阶段都不具竞争力。提示模型能很好地分割比赛,但一个规模小得多的训练模型在识别接触事件方面以极低的成本击败了它,而运动本身的规则能恢复像素无法提供的回合结果。身份识别是所有自动化方法都难以应对之处:球衣号码是一个静态事实,如果从未可见,时间推理无法恢复,这与体育动作不同——动作是重复的运动模式,时间模型可以利用,这正是整体推理能提升事件检测而身份识别保持不变的原因。最后,我们总结了每种方法在何处值得其成本,以及哪些经验能超越排球,推广到业余体育领域。
英文摘要
Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as evidence a model is robust enough for deployment. Amateur team sport is a useful, largely untested place to check that assumption: over eight million students played a school sport in the United States in 2024-25 alone, almost none of it filmed by more than a single fixed camera, with several candidate actors crowded into frame and no operator or second angle to fall back on. Using volleyball as a test case, we ask whether strong performance on general video and world-model benchmarks translates into reliable, per-player attribution once footage is this chaotic, turning footage into statistics through a chain of tasks from finding play boundaries to naming who did what. We evaluate four approaches (prompting and agentic reasoning over frontier vision-language models, classical computer vision with small trained specialists, self-supervised video world models, and manual annotation) at every stage, on 66 amateur matches with 46,648 human-labelled contacts, filmed under conditions no published benchmark uses. No single paradigm wins every stage, and static, single-frame computer vision is not competitive at any stage involving motion or identity. A prompted model segments matches well, yet a far smaller trained model beats it at spotting contacts for a fraction of the cost, and the sport's own rules recover rally outcomes the pixels cannot. Identity is where every automated approach struggles: a jersey number is a static fact temporal reasoning cannot recover if never visible, unlike sporting action, a repeated motor pattern a temporal model can exploit, which is why holistic reasoning improves event detection while identity stays unchanged. We close with where each approach earns its cost, and what transfers beyond volleyball to amateur sport.