arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07127cs.CVcs.AI

PlaySuite:面向交互式视觉智能的大规模基准

PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence

Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini, Michelle Lorena Acevedo Callejas, Mohammad Mahdi Derakhshani, Kristof Meding, Joaquin Vanschoren, Cees G. M. Snoek

首次发表
浏览论文内容

中文总结 AI 辅助

PlaySuite是一个大规模基准,通过5000多个开源游戏评估交互式视觉智能,揭示当前模型存在感知-行动差距,为动态环境中的智能体发展提供测试平台。

中文摘要 AI 辅助

多模态基础模型的最新进展在静态感知和推理基准上取得了强劲表现,然而此类评估在很大程度上忽略了一个智能的核心方面:在动态环境中长时间跨度内胜任行动。我们推出了PlaySuite,这是一个大规模基准,用于评估跨超过5000个从PyWeek和此HTTP URL精选的开源视频游戏的交互式视觉智能。这些独立游戏涵盖多种类型和引擎,包括Pygame、HTML5、Godot和Unity,对于当前模型而言大多属于分布外数据,从而降低了通过检索记忆中的攻略或网络规模训练产物来取得成功的可能性。为了实现对异构标题的可扩展评估,我们开发了一个针对HPC集群优化的统一闭环交互框架,以及一个视频-LLM作为评判者的协议,该协议将可观察的游戏内里程碑映射到标准化的进度级别。我们评估了十四个近期开放模型,涵盖视觉语言模型、计算机使用智能体和视觉语言行动模型。我们的结果提供了感知-行动差距的有力证据:尽管具备强大的推理能力,当前模型在持续进展上挣扎,并在空间定位、行动执行和自我纠正方面表现出反复出现的失败。PlaySuite提供了一个可复现且可扩展的测试平台,用于衡量从视觉感知到目标导向交互的进展,并为开发能够在动态视觉环境中行动、适应和泛化的模型奠定了基础。

英文摘要

Recent advances in multimodal foundation models yield strong performance on static perception and reasoning benchmarks, yet such evaluations largely overlook a central aspect of intelligence: acting competently in dynamic environments over extended time horizons. We introduce PlaySuite, a large-scale benchmark for evaluating interactive visual intelligence across more than 5K open-source video games curated from PyWeek and itch.io. Spanning diverse genres and engines, including Pygame, HTML5, Godot, and Unity, these independent games are largely out-of-distribution for current models, reducing the likelihood that success can be achieved by retrieving memorized walkthroughs or web-scale training artifacts. To enable scalable evaluation across heterogeneous titles, we develop a unified closed-loop interaction framework optimized for HPC clusters alongside a Video-LLM-as-a-judge protocol that maps observable gameplay milestones to standardized progress levels. We evaluate fourteen recent open models spanning vision-language models, computer-use agents, and vision-language-action models. Our results yield strong evidence of a perception-action gap: despite strong reasoning capabilities, current models struggle to make sustained progress and exhibit recurring failures in spatial grounding, action execution, and self-correction. PlaySuite provides a reproducible and extensible testbed for measuring progress from visual perception to goal-directed interaction, and a foundation for developing models that can act, adapt, and generalize in dynamic visual environments.

↑