GameHorizon Suite:游戏玩法中的多时域数据与评估
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
- ARC Lab, Tencent(腾讯ARC实验室)
- GVC Lab, Great Bay University(大湾区大学GVC实验室)
- University of Macau(澳门大学)
- National University of Singapore(新加坡国立大学)
- Huazhong University of Science and Technology(华中科技大学)
- MMLab, CUHK(香港中文大学多媒体实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有游戏AI评估缺乏多时域指令和大规模数据的问题,提出GameHorizon Suite,包含自动化标注、5,000小时21款游戏数据集及离线在线测试,评估47个模型,提供标准化评估尺度。
AI中文摘要:
现代视频游戏为AI模型提供了一个可测量的测试平台,结合了视觉理解、指令分解、目标规划以及在多个时间跨度上的精确动作控制能力。然而,现有的数据集和基准要么覆盖的游戏范围狭窄,要么缺乏语言指令,要么依赖于高方差的在线回滚。为解决这些挑战,我们引入了GameHorizon,一个统一的数据与评估套件,用于衡量不同模型家族在不同时间跨度上的游戏玩法能力。GameHorizon Suite由三个组件构成。首先,GameHorizon-Annotator是一个可扩展且自动化的多时域指令标注流水线。其次,利用该流水线,我们构建了GameHorizon-Data,这是首个大规模AAA游戏玩法数据集,包含时间对齐的视频、玩家动作和多时域指令。它包含来自21款游戏的5,000小时录制内容,由100名人类专家玩家收集。第三,我们构建了GameHorizon-Bench,提供可复现的离线测试和逐步在线测试。离线轨道通过数千个标准化问题实现可复现的评估,这些问题组织为三个主要任务和一系列诊断变体,而在线轨道则测试离线分数是否反映实际游戏玩法能力,并将失败定位到长时域游戏玩法中的特定步骤。基于我们的GameHorizon Suite,我们通过超过一百万次模型调用评估了47个模型,揭示了任务难度的有意义层级和模型能力的显著差异。我们的工作可以为跨时间跨度和模型家族的游戏玩法能力评估提供标准化尺度。我们将发布我们的数据集、标注器和基准,以促进未来研究。
英文摘要:
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.