发表机构
Department of Computer Science and Engineering(计算机科学与工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究利用大语言模型生成多场景游戏世界时面临的问题,提出MAGIC系统,通过四阶段管道将自然语言提示转化为可运行项目,解决跨场景一致性等问题,在新基准测试中表现出色,生成可执行项目且过渡识别指标优。
AI 中文摘要
多场景导航是当代3D游戏的一个关键特性,但创作起来很费力。近期大语言模型和多模态大语言模型场景生成器使单场景合成成本大幅降低,但无法生成连接的多场景世界。我们识别出单场景方法未解决的三个障碍:跨场景一致性、场景内可导航性以及过渡是否可行的评估。我们提出了MAGIC,一个能解决所有这三个问题从提示到项目的系统。它是一个四阶段管道,将单个自然语言提示转化为可运行的多场景游戏项目。因为现有单场景保真度指标从不执行过渡,我们还引入了一个专注于过渡的评估代理。在100个多场景案例的新基准测试中,MAGIC为每个案例生成可执行项目,在端到端过渡识别上达到0.99精度、0.95召回率和0.96 F1。逐阶段来看,它比大语言模型基线和Holodeck能恢复更多真实门户并产生明显更具可导航性的布局。
英文摘要
Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a defining feature of contemporary 3D games, but authoring it is laborious: every portal must have consistent endpoints on both sides, each interior must remain navigable once it is furnished, and the resulting connectivity must be kept consistent across many files. Recent large language model (LLM) and multimodal LLM (MLLM) scene generators have made single-interior synthesis dramatically cheaper, yet they produce one scene at a time and cannot, by naive repetition, yield a connected multi-scene world. We identify three obstacles that single-scene methods leave unsolved: cross-scene consistency, in-scene navigability, and the evaluation of whether a transition actually works. We present MAGIC, a prompt-to-project system that addresses all three. MAGIC is a four-stage pipeline that turns a single natural-language prompt into a runnable multi-scene game project: it plans a shared transition-aware intermediate representation, specifies each scene while enforcing portal reachability with a flood-fill validator, generates the scenes together with their transition scripts, and combines them into one project. Because existing single-scene fidelity metrics never execute a transition, we further introduce a transition-focused evaluation agent that runs each transition in play. On a new benchmark of 100 multi-scene cases, MAGIC produces an executable project for every case and reaches 0.99 precision, 0.95 recall, and 0.96 F1 on end-to-end transition identification; stage by stage, it recovers more ground-truth portals and yields markedly more navigable layouts than an LLM baseline and Holodeck. Our code is available at https://github.com/sereneee1201/MAGIC/.