arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PlayWorld:基于智能体玩家的长程目标世界模型基准测试

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao

arXiv 2608.13552首次发表:更新:

发表机构

The Chinese University of Hong Kong; The University of Hong Kong; Zhejiang University; Kuaishou Technology(香港中文大学; 香港大学; 浙江大学; 快手科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有世界模型跨模型公平比较的难题,推出含171个场景的PlayWorld基准,通过多模态智能体玩家从多维度评估9种先进世界模型,发现其长程交互式目标表现仍不可靠。

AI 中文摘要

视频世界模型会基于当前观测和用户动作模拟未来状态。近期的系统已在长序列上展现出令人印象深刻的视频一致性与动作可控性,但对这些交互式模型进行公平比较仍颇具挑战。实际中,人类玩家通常通过交互追求长程目标来评估世界模型,例如用户可能会转动360度查看环境是否保持一致,或走入水中检查是否生成真实的水波纹,而达成同一目标所需的动作序列在不同模型间可能差异巨大,这使得固定动作条件下的评估不适用于跨模型比较。为解决该问题,我们采用多模态智能体玩家与世界模型交互以达成指定的长程目标,基于此范式引入PlayWorld基准,提供171个场景,每个场景均有指定目标。为全面评估性能,我们从四个核心维度进行评估:几何一致性、交互保真度、视线外演化、洞察演化,此外还纳入视频质量与可控性的基础能力指标。对9种最先进世界模型的实验表明,当前模型在长程交互式目标上仍不可靠,尤其在维持空间一致性与持久状态演化方面表现不足。代码与数据可在该httpsURL获取。

英文摘要

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.

Commentsproject page: https://kxding.github.io/project/PlayWorld/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑