Text2Villa:通过物理感知的合成分析进行3D室内环境的分层生成
Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis
浏览论文内容
中文总结 AI 辅助
研究旨在解决从自然语言生成3D室内场景的问题,提出Text2Villa框架。宏观构建数据集微调布局生成器,微观引入A-PSSG并结合碰撞检测与语义推理,有效解决相关问题,性能优于以往方法,为下游应用提供3D内容基础。
中文摘要 AI 辅助
从自然语言生成3D室内场景潜力巨大,但现有方法大多无法生成具有垂直连通性和任意多边形边界的多房间结构,且缺乏对连续3D物理定律的深入理解,导致严重的几何穿透和漂浮伪影。本文提出Text2Villa这一新颖的分层生成框架。宏观上构建多层数据集微调自回归布局生成器;微观上引入A-PSSG将物理可供性抽象为节点属性并建立约束。通过集成几何碰撞检测引擎与多模态大语言模型的高级语义推理,求解器有效解决碰撞等问题。实验表明Text2Villa性能优于以往方法,能从文本生成高保真且符合物理原理的别墅级3D环境,为下游应用提供可靠基础。
英文摘要
Generating 3D indoor scenes from natural language holds tremendous potential, yet existing methods predominantly fail to generate multi-room structures with vertical connectivity and arbitrary polygonal boundaries. Furthermore, they lack a deep grounding in continuous 3D physical laws, leading to severe geometric penetrations and floating artifacts. In this work, we propose Text2Villa, a novel hierarchical generative framework. At the macro level, we construct a multi-story dataset to fine-tune an autoregressive layout generator, ensuring the direct parsing of text into 3D building foundations featuring polygonal boundaries and multi-story connectivity. To enforce physical laws during micro-level asset arrangement, we introduce the Affordance-driven Physical-Semantic Scene Graph (A-PSSG) to explicitly abstract physical affordances (such as support surfaces and containment cavities) into node attributes, establishing strict geometric and semantic edge constraints. Guided by the A-PSSG, we formulate scene instantiation as a constrained closed-loop optimization problem following the analysis-by-synthesis paradigm. By integrating an underlying geometric collision detection engine with the high-level semantic reasoning of multimodal large language models (MLLMs), our heuristic solver dynamically executes physics-aware actions under the observation-evaluation-modification mechanism to effectively resolve mesh collisions, floating artifacts, and fine-grained cavity containment failures. Extensive experiments demonstrate that Text2Villa outperforms previous methods across various metrics, robustly generating high-fidelity and physically plausible villa-level 3D environments from text, thereby providing a reliable and interactive 3D content foundation for downstream applications.