arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05416cs.CV

WorldSculpt:基于接地视频生成组合式世界

WorldSculpt: Generating Compositional Worlds from Grounded Videos

  • Alaya Lab(阿里达摩院Alaya实验室)
  • The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

Muyao Niu, Jixuan He, Ruihan Yu, Lian Fu, Yonghao Yu, Zheng-Hui Huang, Yifan Zhan, Fengbo Lan, Yongtao Ge, Yinqiang Zheng, Kaipeng Zhang, Zhixiang Wang

中文总结 AI 辅助

本研究提出WorldSculpt,通过适配单物体3D生成先验到多视图观测,实现从接地视频生成含数百物体的复杂组合式3D世界,所提Pixal3D方法在基准测试中优于现有技术,还可转换3DGS世界为组合网格场景。

中文摘要 AI 辅助

我们研究的问题是生成包含数百个物体的杂乱场景的组合式3D表示,目标是将该场景表示为放置在共享世界坐标系中的单个物体网格的集合,以满足游戏、AR/VR、仿真和机器人技术等下游应用的需求。在密集杂乱的场景中,该任务极具挑战性,因为物体之间存在严重的相互遮挡,每个视图仅能显示其几何结构的一部分。基于几何的方法通常将场景重建为单一表示,会在遮挡区域留下不完整的几何结构,而现有的带有生成先验的组合方法大多仅限于相对简单的场景。我们证明,通过将强大的单物体3D生成先验适配到多视图观测中,就可以组合式地生成包含数百个物体的复杂场景。我们将该范式实例化为Pixal3D,为其扩展了多视图条件通路,该通路将物体生成建立在多个带姿态的观测基础上。尽管该模型完全是在规范空间中的单物体上进行微调的,但它无需任何场景级别的训练就能泛化到存在严重遮挡的大型场景,证明了该范式的可行性和可扩展性。我们还引入了UE-MeshyScene,这是一个包含数百个物体的密集杂乱场景的 photorealistic(照片级真实感)基准,带有逐物体注释和真实网格。在单物体、受控多物体和UE-MeshyScene评估中,我们的方法始终优于现有方法,且随着场景复杂度和遮挡程度的增加,性能提升更为显著。最后,我们通过将生成的3DGS世界(如Marble和HY-World 2.0)转换为组合式网格场景,展示了其更广泛的适用性。

英文摘要

We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally by adapting a strong single-object 3D generative prior to multi-view observations. We instantiate this paradigm with Pixal3D, extending it with a multi-view conditioning pathway that grounds object generation in multiple posed observations. Although the model is finetuned entirely on single objects in canonical space, it generalizes to large scenes with severe occlusion without any scene-level training, demonstrating the feasibility and scalability of this paradigm. We further introduce UE-MeshyScene, a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes. Across single-object, controlled multi-object, and UE-MeshyScene evaluations, our method consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase. Finally, we demonstrate broader applicability by converting generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes.

补充信息

相关深度报道

↑