AI 中文总结
针对现有篮球理解基准仅单一评估能力的不足,构建多模态基准BasketballBench,并提出智能体BasketballSkills,实验显示其在整合多能力的篮球理解任务上优于现有MLLM。
AI 中文摘要
理解篮球比赛需要识别事件、定位动作、识别球员,并将这些与结构化的比赛知识关联起来。现有基准主要一次评估这些能力中的一项,导致这些能力之间的交互未得到充分探索。我们推出BasketballBench,这是一个多模态基准,包含文本、图像和视频领域的10项任务共7980个问题,基于2025-2026赛季NBA构建,包含官方逐场比赛数据、530名现役球员的名单和资料,以及2501个控球级别的转播片段。我们进一步提出BasketballSkills,这是一个智能体,它由8种篮球专用感知与检索工具组成,这些工具属于4种可复用技能,这些技能明确了工具顺序、证据绑定和停止条件。实验表明,当前多模态大语言模型(MLLM)在需要整合多种能力的问题上表现尤其吃力,而BasketballSkills的表现优于这些模型,凸显了明确组合领域专用能力对实现全面篮球理解的有效性。
英文摘要
Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game knowledge. Existing benchmarks primarily evaluate these abilities one at a time, leaving the interactions among these abilities under-explored. We introduce BasketballBench, a multimodal benchmark comprising 7,980 questions across ten tasks in text, image, and video. It is built from the 2025-2026 NBA season and includes official playby-play, rosters and profiles for 530 active players, and 2,501 possession-level broadcast clips. We further propose BasketballSkills, an agent that composes eight basketball-specific perception and retrieval tools under four reusable skills that specify tool order, evidence bindings, and stopping conditions. Experiments show that current MLLMs struggle particularly on questions requiring the integration of multiple capabilities, whereas BasketballSkills outperforms them, highlighting the effectiveness of explicitly composing domain-specific capabilities for comprehensive basketball understanding.
Comments26 pages, 3 figures