发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍 ABot-3DWorld 0 通用多模态 3D 世界模型,核心是 SGP,通过特定流程将多模态输入转换为 3D 世界,用于通用 3D 内容创建,在开源方法中领先,多模态输入下场景保真度更高。
AI 中文摘要
我们展示了 ABot-3DWorld 0,这是一个通用的多模态 3D 世界模型,可将文本、图像和视频输入转换为高保真、可探索的 3D 世界。我们框架的核心是统一的空间生成原语(SGP),它由高质量全景图和空间点云组成,能高效描述任何 3D 空间。多模态输入先提升到该原语中,然后 3D 一致的全景视频生成器沿规划轨迹探索,最后全景视频重建引擎将生成的视频转换为清晰、逼真的 3D 高斯溅射(3DGS)世界。该流程涵盖两种情况:丰富输入通过严格几何恢复提升到 SGP,单张图像或句子通过生成完成。结果是一个用于通用 3D 内容创建的低门槛引擎,实验表明它在开源方法中处于领先地位,在丰富多模态输入下比 Marble 具有更高的场景保真度。
英文摘要
We present ABot-3DWorld 0, a universal multimodal 3D world model that turns text, image, and video inputs into high-fidelity, explorable 3D worlds. At the heart of our framework is a unified Spatial Generative Primitive (SGP), a compact tuple of a high-quality panorama and a spatial point cloud that delivers an efficient description of any 3D space. Multimodal inputs are first lifted into this primitive; a 3D-consistent panoramic video generator then explores the primitive along a planned trajectory; finally, our panoramic video reconstruction engine converts the generated video into a clean, photorealistic 3D Gaussian Splatting (3DGS) world. This pipeline covers two regimes: rich inputs (multi-view sets, casual video) are lifted into the SGP through a geometry-rigorous recovery that mirrors the observed scene, while a single image or sentence is completed generatively into a creative world. The result is one low-barrier engine for general 3D content creation that further anchors generated worlds to geographic points of interest, enabling map-native spatial exploration at consumer scale. Experiments show that ABot-3DWorld 0 sets the state of the art among open-source methods and demonstrates stronger scene fidelity than Marble under rich multimodal inputs.
CommentsOfficial Page: https://abot-world.amap.com/plaza