世界编辑:在日益加深的可执行世界中进行干预
World Editing: Intervening on Executable Worlds at Increasing Depth
浏览论文内容
中文总结 AI 辅助
本研究提出世界编辑概念,通过干预深度描述编辑强度,并构建IGMBench基准(110任务,覆盖Minecraft和Terraria),发现前沿编码智能体在任务级解决78.2%,但可靠性随干预深度下降,视觉一致性为独立弱点。
中文摘要 AI 辅助
交互式世界模型日益能够生成环境并在其中行动,然而,对现有可执行世界进行刻意编辑仍是一个未被充分探索的领域。我们将世界编辑定义为在保持应不变属性的同时,对现有世界进行干预,并引入干预深度这一轴,用以描述编辑耦合世界实体、动力学和系统的强度。我们通过行业级游戏模组实例化这一能力,并引入IGMWorld,连同IGMBench,一个包含110个任务和超过1.1K个可执行状态与行为标准的基准,覆盖Minecraft和Terraria。任务涵盖属性、实体、动力学和系统干预,并通过确定性可执行性、行为、保持性和视觉检查进行评估。前沿编码智能体已展现出可观的世界编辑能力:最强配置在严格的任务级标准下解决了78.2%的任务,而标准级性能达到94.8%。可靠性通常随干预深度增加而降低,这一模式在具有相似评估标准数量的任务中依然存在。大多数失败的编辑仍能成功构建和加载,表明主要难点在于使编辑后的世界按预期行为运行。视觉一致性仍是独立的弱点,所有评估配置的联合视觉通过率均低于50%。这些结果表明,世界编辑是区别于世界生成和交互的一种独特能力,且可执行游戏为研究这一能力提供了实用的测试平台。
英文摘要
Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through deterministic executability, behavioral, preservation, and visual checks. Frontier coding agents already exhibit substantial world-editing capability: the strongest configuration solves 78.2% of tasks under a strict task-level criterion, while criterion-level performance reaches 94.8%. Reliability generally decreases with intervention depth, and this pattern persists even among tasks with similar numbers of evaluation criteria. Most failed edits still build and load successfully, suggesting that the main difficulty is making the edited world behave as requested. Visual consistency remains a separate weakness, with all evaluated configurations below 50% joint visual pass rate. These results show that world editing is a distinct capability from world generation and interaction, and that executable games provide a practical testbed for studying it.
发表机构
- University of Waterloo(滑铁卢大学)
- Comfy Org Research(Comfy Org 研究院)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。