arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05248cs.AIcs.CV

WorldClaw:大规模智能体驱动的3D开放世界生成

WorldClaw: Agentic 3D Open-World Generation at Scale

Chunchao Guo, Jinpeng Li, Yang Li, Zilong Huang

首次发表
浏览论文内容

中文总结 AI 辅助

WorldClaw是一种完全智能体驱动的由粗到细框架,可从开放式文本生成兼具全局空间一致性、丰富局部内容与可编辑实例级资产的大规模3D开放世界场景。

中文摘要 AI 辅助

从开放式文本生成可自由探索的大规模3D世界仍具挑战性,因为系统需同时维持全局空间一致性、丰富的局部内容及适合下游编辑与复用的显式资产。我们提出WorldClaw,这是一种完全智能体驱动、由粗到细的开放世界3D场景生成框架。规划智能体将文本提示转换为区域、地形、资产、材质及空间关系的结构化规范。WorldClaw随后从语义布局、可复用资产、生成式或程序性材质,以及感知区域的高度场构建全局一致的地形基础。对于细节要求高的区域,它生成地形条件下的合成内容,重建可编辑的带纹理网格并恢复其在地形上的放置;基于渲染的智能体进一步优化地形、物体、外观及接触关系。针对多样的开放世界提示,WorldClaw生成的大规模场景具备一致的空间组织、视觉吸引力的局部内容及可编辑的实例级资产,同时保持稳定的全局地形结构。

英文摘要

Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.

补充信息

↑