SceneMosaic:通过混合智能体布局演化实现高效且多样化的仿真就绪场景生成
SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution
浏览论文内容
中文总结 AI 辅助
SceneMosaic结合图像先验与VLM智能体演化,通过局部单元独立演化及笛卡尔积组合,实现高效、多样且物理有效的室内场景生成,在SceneEval-100上匹配最强基线并提速24倍。
中文摘要 AI 辅助
多样且仿真就绪的室内场景对于交互式娱乐和具身人工智能至关重要,然而其可扩展生成仍然具有挑战性。最近依赖视觉语言模型(VLM)的智能体文本到3D场景流程能够生成高保真度的场景,但需要昂贵的迭代式物体放置和细化。另一种主流范式,参数化图像到3D场景模型,能够从2D图像学习的强先验中高效生成场景,但往往导致不精确且物理上无效的场景。更重要的是,这两种范式都难以针对单一输入输出多样化的场景,使其难以反映真实场景的动态变化特性。在本文中,我们提出了SceneMosaic,一个结合了两种范式优点的框架。它从学习到的基于图像的先验中获得初始候选,随后通过VLM智能体演化结果,确保效率和物理有效性。在演化过程中,SceneMosaic利用自然场景的局部性,将场景分解为独立的局部单元,允许在每个单元内进行独立演化,然后通过笛卡尔积组合全局场景。在SceneEval-100上,SceneMosaic在语义布局质量上与最强的智能体基线相当,同时实现了24倍加速,大幅减少了物理违规,并获得了最高的人工评分。我们的代码在此https URL公开可用。
英文摘要
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet their scalable generation remains challenging. Recent agentic text-to-3D scene pipelines that rely on vision-language models (VLMs) can generate scenes of high fidelity but require costly iterative object placement and refinement. Another mainstream paradigm, parametric image-to-3D scene models, produces scenes efficiently from strong priors learned from 2D images but often leads to imprecise and physically invalid scenes. More importantly, both paradigms struggle to output diverse scenes for a single input, making it hard for them to reflect the dynamically changing nature of real scenes. In this paper we propose \textbf{SceneMosaic}, a framework that combines the merits of both paradigms. It obtains the initial candidate from the learned image-based prior, and subsequently evolves the result through VLM agents, ensuring both efficiency and physical validity. Within the evolution process, SceneMosaic exploits the locality of natural scenes and decomposes a scene into independent local units, allowing separate evolution within each unit before composing the global scene via Cartesian product. On SceneEval-100, SceneMosaic matches the strongest agentic baseline in semantic layout quality with a 24x speedup, substantially reduces physical violations, and receives the highest human ratings. Our code is publicly available at https://github.com/rxjfighting/SceneMosaic.
发表机构
- The University of Hong Kong(香港大学)
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。