发表机构
School of Information Science and Technology, Beijing University of Technology; Fudan University; Institute of Automation, Chinese Academy of Sciences(北京工业大学信息科学与技术学院; 复旦大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有篮球评论生成方法与数据集不适用于连续流的局限,推出大规模基准NBA_Streaming,提出因果两阶段框架,可有效提升在线细粒度篮球评论生成性能。
AI 中文摘要
实时篮球评论生成需要确定事件何时足够可观测,并在后续事件展开前对其进行描述。然而,现有方法主要针对预分割片段或完整视频设计,不适合连续流场景;现有数据集在球员身份、细粒度动作、事件属性及连贯事件链方面的监督信息有限,限制了生成评论的事实丰富性。为解决这些局限,我们推出NBA_Streaming——一个用于在线细粒度篮球评论生成的大规模基准测试,包含307小时篮球转播内容和约35K个时间对齐的事件,标注了事件边界、球员身份、细粒度动作、事件链及自然语言评论。从孤立片段转向连续流的设计,使NBA_Streaming能在因果约束下统一评估事件定位、响应可靠性、事实依据及评论质量。我们进一步提出因果两阶段框架,结合先完成后定位的策略与以球为中心的语义 grounding,使系统能从观测流中识别完整事件,并组织场景、事件、身份及动作线索以生成评论。大量实验表明NBA_Streaming具有挑战性,现有基线在在线时序、事实依据及细粒度描述上表现不佳;我们的框架持续优于强替代方案,而剩余差距凸显了NBA_Streaming作为流媒体体育视频理解与生成领域有价值基准的意义。
英文摘要
Live basketball commentary generation requires determining when an event is sufficiently observable and describing it before subsequent events unfold. However, existing methods are primarily designed for pre-segmented clips or complete videos, making them unsuitable for continuous streams. Existing datasets also provide limited supervision for player identities, fine-grained actions, event attributes, and coherent event chains, restricting the factual richness of generated commentary. To address these limitations, we introduce NBA_Streaming, a large-scale benchmark for online fine-grained basketball commentary generation. It contains 307.5 hours of basketball broadcasts and approximately 35K temporally aligned events, with annotations of event boundaries, player identities, fine-grained actions, event chains, and natural-language commentary. By moving from isolated clips to continuous streams, NBA_Streaming enables unified evaluation of event localization, response reliability, factual grounding, and commentary quality under causal constraints. We further propose a causal two-stage framework that combines completion-first localization with ball-centric semantic grounding, enabling the system to identify complete events from observed streams and organize scene, event, identity, and action cues for commentary generation. Extensive experiments reveal the difficulty of NBA_Streaming, where existing baselines struggle with online timing, factual grounding, and fine-grained description. Our framework consistently improves over strong alternatives, while the remaining gap highlights NBA_Streaming as a valuable benchmark for streaming sports video understanding and generation. The dataset is publicly available at https://github.com/wyy-081/NBA_Streaming.