arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SportsGrounder:用于密集体育视频推理的提案辅助交错定位框架

SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning

Yizhi Li, Jiawei Jiang, Guanhong Wang, Yingcai Wu, Gaoang Wang

arXiv 2608.07932首次发表:更新:

发表机构

Zhejiang University; Beijing Jiaotong University(浙江大学; 北京交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LMMs难以处理密集体育视频细粒度推理的问题,提出SportsGrounder框架,结合IGF机制与AAS模块,在SoccerNet等数据集上实现最优准确率。

AI 中文摘要

体育视频分析对于运动分析和转播优化至关重要。然而,密集体育视频推理需要在长时序上下文下对大量小规模、高度交互且视觉同质化的实体(例如身着相同队服的球员、球)进行细粒度理解。当前的大型多模态模型(LMMs)在处理这类密集视觉复杂性时存在固有缺陷,由于缺乏细粒度视觉细节,这些模型往往过度依赖文本先验来猜测答案,尤其是在区分视觉相似的动作和球员时。为解决该问题,我们提出SportsGrounder,这一框架利用开放词汇视觉专家辅助专门针对密集体育视频推理的交错定位。为实现精确的空间定位,我们提取领域引导的对象提案,并引入交错定位融合(IGF)机制,该机制逐帧将显式边界框坐标、隐式视觉语义与全局网格特征相融合,此设计保留了严格的时序对齐并避免序列长度爆炸。此外,我们设计了动作感知监督(AAS)模块,直接正则化模型的隐藏状态,迫使网络学习准确的运动表征而非依赖语言偏差。通过混合偏好优化(MPO)进行优化以更好地区分具有迷惑性的干扰项,我们在新整理的密集体育VQA数据集(源自SoccerNet和FineSports)上开展的大量实验表明,SportsGrounder显著提升了细粒度推理能力并达到了当前最优的准确率。

英文摘要

Sports video analysis is crucial for athletic analytics and broadcasting enhancement. Dense sports video reasoning, however, demands a fine-grained understanding of numerous small-scale, highly interactive, and visually homogeneous entities (e.g., players sharing identical uniforms, the ball) across long temporal contexts. Current Large Multimodal Models (LMMs) inherently struggle with such dense visual complexities. Due to the lack of fine-grained visual details, these models often over-rely on textual priors to guess answers, especially when distinguishing visually similar actions and players. To address this, we propose SportsGrounder, a framework that leverages an open-vocabulary visual expert to aid interleaved grounding specifically for dense sports video reasoning. To achieve precise spatial localization, we extract domain-guided object proposals and introduce an Interleaved Grounding Fusion (IGF) mechanism. The IGF frame-by-frame integrates explicit bounding box coordinates and implicit visual semantics with global grid features. This design preserves strict temporal alignment and prevents sequence length explosion. Furthermore, we design an Action-Aware Supervision (AAS) module that directly regularizes the model's hidden states, forcing the network to learn accurate motion representations rather than relying on language bias. Optimized with Mixed Preference Optimization (MPO) to better distinguish deceptive distractors, our extensive experiments on newly curated dense sports VQA datasets (derived from SoccerNet and FineSports) demonstrate that SportsGrounder significantly improves fine-grained reasoning and achieves state-of-the-art accuracy.

CommentsACMMM 2026

DOI:10.1145/3767308.3836263

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑