SVI-Bench: 一个用于战略视频智能的动态微世界
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
- UNC Chapel Hill(北卡罗来纳大学教堂山分校)
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出SVI-Bench,一个基于团队体育动态微世界的大规模基准,通过四个层级(动态场景理解、因果推理、战略模拟、智能体合成)的9个任务评估视频智能从感知到战略规划的能力,发现模型在感知任务上表现良好但在认知层级上性能急剧下降。
AI中文摘要:
真正的视频智能需要的不仅仅是识别可见内容:它需要推理事件为何发生,预测在不同条件下会有什么变化,并决定下一步该做什么。我们将这种从感知到因果推理、模拟再到战略规划的演进称为战略视频智能(SVI)。现有基准均未评估这一能力栈:野外视频缺乏因果和战略问题的可验证真实数据,而合成环境则牺牲了真实多智能体系统的复杂性。为弥补这一差距,我们引入了SVI-Bench,这是一个大规模基准,利用团队体育作为动态微世界,将真实世界多智能体交互(10-22个智能体在对抗压力下做出协调决策)的复杂性与显式规则和确定性结果的可验证性相结合。SVI-Bench包含约3.5万小时的广播视频、1500万个标注动作、1.5万小时的专家解说、2.3万份比赛报告以及涵盖篮球、足球和冰球的10.3万条结构化统计记录,所有这些均通过一个将原始比赛数据转换为密集交叉引用语料库的数据引擎构建。我们将评估组织为9个任务,涵盖一个渐进的四层层次结构:动态场景理解、因果推理、战略模拟和智能体合成。评估强多模态和智能体基线后,我们发现一个能力悬崖:模型在感知任务上表现胜任,在细粒度动作问答上达到约73%的准确率,但在每个后续认知层级上急剧下降。智能体任务最为困难:当需要自主收集并整合来自180万个片段语料库的证据时,最强模型仅达到5%的准确率。
英文摘要:
True video intelligence demands more than recognizing what is visible: it requires reasoning about why events unfold, predicting what would change under different conditions, and deciding what to do next. We refer to this progression, from perception through causal reasoning and simulation to strategic planning, as Strategic Video Intelligence (SVI). No existing benchmark evaluates this capability stack: in-the-wild videos lack verifiable ground truth for causal and strategic questions, while synthetic environments sacrifice the complexity of real multi-agent systems. To bridge this gap, we introduce SVI-Bench, a large-scale benchmark that leverages team sports as a dynamic microworld, combining the complexity of real-world multi-agent interaction (10-22 agents making coordinated decisions under adversarial pressure) with the verifiability of explicit rules and definitive outcomes. SVI-Bench comprises approximately 35K hours of broadcast video, 15M annotated actions, 15K hours of expert commentary, 23K game reports, and 103K structured statistical records across basketball, soccer, and hockey, all constructed via a data engine that transforms raw game data into a dense, cross-referenced corpus. We organize evaluation into 9 tasks spanning a progressive four-pillar hierarchy: Dynamic Scene Understanding, Causal Reasoning, Strategic Simulation, and Agentic Synthesis. Evaluating strong multimodal and agentic baselines, we find a capability cliff: models perform competently on perceptual tasks, achieving approximately 74% on fine-grained action QA, but degrade sharply at each successive cognitive level. Agentic tasks prove hardest: the strongest model achieves only 5% accuracy when required to autonomously gather and integrate evidence across a corpus of 1.8M clips.