AI 中文总结
该研究针对VLM驱动3D室内布局生成中全局一致性与物理可行性不足的问题,提出结合GSV与GPFS的混合策略,实现了该任务的当前最优性能。
AI 中文摘要
我们研究开放词汇3D室内布局生成任务,该任务利用无标注3D资产结合自由形式语言指令合成多样化且符合物理规律的场景。近期方法借助大语言模型(LLMs)和视觉语言模型(VLMs)从文本生成结构化场景,但多数模型隐式处理资产间关系,或依赖局部成对约束与局部优化,这些公式化方法与全局、高度非凸的布局空间适配性差,常生成局部合理但全局不一致或物理不可行的场景。我们提出一种基于图的中间表示,将语义一致性与物理可行性分离,结合混合搜索-优化策略解决该问题。首先,全局语义验证(GSV)将场景表示为结构化图,通过基于规则的验证强制执行语义约束,此显式验证可移除矛盾配置,生成全局一致的语义骨架。其次,全局物理可行性搜索(GPFS)结合进化搜索进行全局探索与基于梯度的优化进行局部利用,减少对VLM提议初始化的依赖,提升在非凸且不连续可行空间中的鲁棒性。GSV与GPFS共同推动布局生成从局部关系建模和初始化敏感型优化转向全局一致的推理与搜索。实验表明,我们的方法在开放词汇3D室内布局生成任务中达到了当前最优性能,同时提升了语义一致性与物理可行性。
英文摘要
We study open-vocabulary 3D indoor layout generation, which synthesizes diverse and physically plausible scenes from unlabeled 3D assets using free-form language instructions. Recent methods leverage large language models (LLMs) and vision-language models (VLMs) to generate structured scenes from text. However, most model inter-asset relations implicitly or rely on local pairwise constraints and local optimization. These formulations are poorly aligned with the global, highly non-convex layout space, often yielding locally plausible yet globally inconsistent or physically infeasible scenes. We address this problem with a graph-based intermediate representation that separates semantic coherence from physical feasibility, together with a hybrid search-and-refinement strategy. First, Global Semantic Verification (GSV) represents scenes as structured graphs and enforces semantic constraints through rule-based verification. This explicit validation removes contradictory configurations and produces a globally consistent semantic scaffold. Second, Global Physical Feasibility Search (GPFS) combines evolutionary search for global exploration with gradient-based refinement for local exploitation. It reduces dependence on VLM-proposed initialization and improves robustness in non-convex and discontinuous feasible spaces. Together, GSV and GPFS move layout generation beyond local relational modeling and initialization-sensitive optimization toward globally consistent reasoning and search. Experiments show that our method achieves state-of-the-art performance in open-vocabulary 3D indoor layout generation, improving both semantic consistency and physical plausibility.
Comments15 pages, 8 figures