arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TennisVAR:一种基于击球证据的多模态大语言模型,用于网球视频中的战术推理

TennisVAR: A Stroke-Evidence-Grounded Multimodal Large Language Model for Tactical Reasoning in Tennis Videos

Yifan Mei, Qingling Shi, Changli Wu, Jiayuan Rao, Jiayi Ji, Liujuan Cao

arXiv 2608.12920首次发表:更新:

AI 中文总结

该研究针对网球视频理解的感知与理解差距,提出回合级战术推理任务,构建专家标注基准TRACE,开发基于证据的多模态大语言模型TennisVAR,实现网球视频的战术推理。

AI 中文摘要

体育视频理解正从事件识别向解释动作如何共同塑造比赛进程发展,但现有网球视频方法要么仅感知单个击球而不建模其战术依赖关系,要么生成高级分析却未将其与底层事件关联。为弥合感知到理解的差距,我们提出基于击球证据的战术推理这一新的回合级任务,要求模型联合预测开放式答案、分层战术标签、有序的支持击球序列及决定性关键动作,且每个证据击球都锚定至其球拍-球接触帧。我们进一步引入TRACE(网球中基于动作链证据的战术推理),这是一个大规模专家标注基准,包含11189个回合视频、41485个击球事件、25429个战术单元及11189个问答对,统一了细粒度击球属性、跨击球战术关系、分层战术标注及覆盖事实感知、战术理解与决策推理的基于证据的问题。基于TRACE,我们提出TennisVAR(网球视频动作链推理器),这是一种基于证据的多模态大语言模型,遵循“事件-关系-证据-战术”推理范式,其中事件解析模块将连续回合转换为明确的击球事件序列,而战术图引导的时间推理器则联合建模回合进程与同球员决策依赖关系,以识别问题相关证据与决定性动作。

英文摘要

Sports-video understanding is moving beyond event recognition toward explaining how actions collectively shape match progression, however, existing tennis-video methods either perceive individual strokes without modeling their tactical dependencies or generate high-level analyses without grounding them in the underlying events. To bridge this perception-to-understanding gap, we formulate stroke-evidence-grounded tactical reasoning, a new rally-level task that requires models to jointly predict an open-ended answer, a hierarchical tactic label, an ordered sequence of supporting strokes, and decisive key actions, with each evidence stroke anchored to its racket-ball contact frame. We further introduce TRACE (Tactical Reasoning with Action-Chain Evidence in Tennis), a large-scale expert-annotated benchmark containing 11,189 rally videos, 41,485 stroke events, 25,429 tactical units, and 11,189 question-answer pairs, which unifies fine-grained stroke attributes, cross-stroke tactical relations, hierarchical tactic annotations, and evidence-grounded questions across factual perception, tactical understanding, and decision reasoning. Building on TRACE, we propose TennisVAR (Tennis Video Action-chain Reasoner), an evidence-grounded multimodal large language model that follows an "event-relation-evidence-tactic" reasoning paradigm, where an Event Parsing Module converts continuous rallies into explicit stroke-event sequences while a Tactical Graph-Guided Temporal Reasoner jointly models rally progression and same-player decision dependencies to identify question-relevant evidence and decisive actions.

CommentsProject Page: https://whynotgit2025.github.io/TennisVAR/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑