发表机构
Zhejiang University; Microsoft Research Asia(浙江大学; 微软亚洲研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对体育视频多视角理解缺乏评估基准及MLLMs难以利用多视角信息的问题,引入SportMV - Bench基准,分析瓶颈所在,并提出SportMV - Agent框架,通过迭代循环实现主动视角选择等,相比最强MLLM基线有显著提升。
AI 中文摘要
近期多模态大语言模型(MLLMs)在单视角视频理解基准测试中表现出色。然而,体育视频存在密集遮挡、快速运动和复杂交互,单视角难以解决。实际中体育赛事由多摄像头记录,可为裁判提供补充证据,但尚无基准测试评估MLLMs对多视角体育视频的理解。为此引入SportMV - Bench基准,含787个多视角视频束和2592个问答对。分析表明当前MLLMs难以有效利用多视角信息,瓶颈在于细粒度视觉感知和视角选择。提出SportMV - Agent框架,实现了14.46%的相对提升。
英文摘要
Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid motion, and complex interactions that are difficult to resolve from a single viewpoint. In practice, sports events are recorded from multiple camera angles, providing complementary evidence used by referees. Yet, no existing benchmark evaluates MLLMs on multi-view sports video understanding. To address this gap, we introduce SportMV-Bench, a comprehensive benchmark built from official match recordings, through a dedicated pipeline combining LLM-based generation, MLLM-based verification, and human filtering to ensure quality and consistency. SportMV-Bench containing 1022 multi-view video bundles and 3015 question-answer pairs spanning 10 sports across three categories: Perception-Aware Recognition (PAR), Rule-aware Event Interpretation (REI), and Adjudicative Decision Reasoning (ADR). Our analysis shows that current MLLMs fail to effectively exploit multi-view information, with the bottlenecks lying in fine-grained visual perception and view selection rather than logical reasoning or domain knowledge. We propose SportMV-Agent, an agentic framework that orchestrates an iterative loop of active view selection, perception tool execution, and evidence-grounded reasoning, achieving a significant 15.61% relative improvement over the strongest MLLM baseline.