AI 中文总结
提出世界智能体任务,将世界生成评估从交付转向持续运行,并构建WorldAgent-Benchmark,通过维护和演绎两轨道衡量模型维持世界运行的能力,发现因果检查是各模型最薄弱环节。
AI 中文摘要
世界模型正从生成逼真帧转向生成可玩世界,然而,所交付的世界能否持续运行,在任何地方都未得到测试。现有评估止步于生成、交付或单步转换,且每一步都在不同节点停止。正确的局部状态转换或中间结果并不能保证因果事件流的正确组织。我们提出世界智能体任务,将世界生成的评估点从交付时刻移至后续的持续运行。在该任务中,模型无需生成世界,而是负责维持世界运行,这要求协调事件并推进其后果以约束后续演化。我们在WorldAgent-Benchmark中实例化该任务,包含两个互补轨道。在维护轨道中,模型必须将连续叙事的事件锚定到显式世界状态的正确转换中,同时遵守因果、时间和并发约束。在演绎轨道中,模型必须在部分观测下预测世界将如何演化,并朝目标行动。维护轨道结合了LLM辅助的语义判断与程序化验证和评分,而演绎轨道完全以程序化方式评估。个体判断可依据世界状态和执行日志进行审计,分数可从保存的判断和执行记录中重新计算。在8个模型中,随着预构建结构从世界中移除,分数稳步下降,且因果关系检查是每个模型最薄弱的环节。该基准使世界的持续运行可测量,并区分局部完成与事件组织中的失败。代码和数据集将在该https URL上发布。
英文摘要
World models are moving from generating realistic frames to generating playable worlds, yet whether a delivered world can keep running is not tested anywhere. Existing evaluations stop at generation, at delivery, or at single-step transitions, and each stops at a different point along the way. Correct local state transitions or intermediate outcomes do not guarantee a correctly organized causal event flow. We propose the world agent task, which moves the evaluation point of world generation from the moment of delivery to the continued operation that follows. In this task, a model is not asked to generate a world. It is held responsible for keeping the world running, which requires coordinating events and carrying forward their consequences to constrain subsequent evolution. We instantiate the task in WorldAgent-Benchmark with two complementary tracks. In the maintenance track, the model must ground the events of a continuous narrative into correct transitions of the explicit world state while respecting causal, temporal, and concurrency constraints. In the deduction track, the model must predict how the world will evolve under partial observations and act toward a goal. The maintenance track combines LLM-assisted semantic judgments with programmatic validation and scoring, while the deduction track is evaluated entirely programmatically. Individual judgments are auditable against world states and execution logs, and scores can be recomputed from the saved judgments and execution records. Across 8 models, scores decline steadily as pre-built structure is removed from the world, and causal-relation checking is the weakest component for every model. The benchmark makes the continued operation of a world measurable and distinguishes local completion from failures in event organization. Code and dataset will be released on https://github.com/HCPLab-SYSU/WorldAgent-Benchmark.