发表机构
Alibaba Token Hub, Alibaba Group; State Key Laboratory of General Artificial Intelligence, BIGAI; Tsinghua University; Nanjing University; State Key Laboratory of General Artificial Intelligence, Peking University(阿里巴巴集团,阿里巴巴Token Hub; 北京通用人工智能研究院,通用人工智能全国重点实验室; 清华大学; 南京大学; 北京大学,通用人工智能全国重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HappyWorld-Bench通过六能力框架和三个轨道评估世界模型,发现视频、空间和具身模型在状态一致性与动作响应上仍存在可靠性差距。
AI 中文摘要
评估世界模型需要同时评估它们生成的世界质量,以及这些世界在探索、交互和修改下的一致性与响应性。我们引入了HappyWorld-Bench,一个全面的基准测试,用于评估生成的世界在智能体与之交互时是否保持可靠。我们的设计基于一个包含六种世界能力(W1-W6)的分层能力框架,从生成式构建到统一世界建模,并在三个独立的评估轨道中实例化:视频世界模型、空间世界模型和具身世界模型。HappyWorld-Bench包含1,138个视频提示、300个空间场景和254个具身测试用例。在所有三个轨道中,我们构建并运行HappyWorld-Arena来组织人类A/B比较,并推导出模型级别的Elo评分,这些评分补充了新设计的自动化指标,以捕捉行为正确性。我们在此统一框架下评估了14个视频世界模型、9个空间系统和8个具身候选模型。结果揭示了所有三个轨道中仍然存在的可靠性差距:视频模型在长时间展开和重访期间表现出降低的一致性,空间模型在放置准确率上最多达到70.14%,编辑执行率达到73.33%,而具身模型难以在多步动作中保持状态,并对改变的动作条件和物理规则做出精确响应。这些发现强调了评估世界模型不仅需要关注视觉质量,还需要关注状态一致性以及它们对动作和干预响应的正确性。
英文摘要
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.