发表机构
Taobao & Tmall Group of Alibaba(阿里巴巴淘宝与天猫集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有室内布局生成方法的缺陷,该研究提出基于LLM的LayoutDSL框架,通过领域特定语言动作空间学习布局策略,结合数据集构建与强化学习优化,显著提升了布局的空间合理性与设计逻辑性。
AI 中文摘要
室内场景布局生成是室内设计中一项具有挑战性的任务。现有方法常通过将房间条件简化为粗略的3D边界框来过度简化该任务,且忽略门、窗等结构元素。更根本的是,许多现有方法将空间推理表述为直接的坐标预测,从而将室内布局设计转化为对原始几何参数的连续回归,这阻碍了模型学习智能布局设计的底层推理逻辑。我们提出了LayoutDSL,一种基于大型语言模型(LLM)的新型框架,用于在领域特定语言(DSL)动作空间中学习室内布局策略。该DSL提供了布局信息的显式符号表示,并作为布局推理的结构化动作空间,其中每个动作对应一个可解释的设计决策。在这种基于DSL的策略学习范式下,我们构建了3D-FrontDSL,这是一个房间结构注释与合成DSL动作序列配对的数据集,用于监督微调。为了获得更具泛化性和可扩展性且具有可验证反馈的策略,我们设计了基于室内设计原则和物理合理性的奖励,并通过强化学习优化该策略。大量实验表明,与强大的基线和现有方法相比,LayoutDSL显著提升了空间合理性和设计逻辑性。
英文摘要
Indoor scene layout generation is a challenging task in interior design. Existing methods often oversimplify the task by reducing room conditions to coarse 3D bounding boxes and neglecting structural elements such as doors and windows. More fundamentally, many prior approaches formulate spatial reasoning as direct coordinate prediction, thereby casting interior layout design as continuous regression over raw geometric parameters, which hinders the model from learning the underlying reasoning logic of intelligent layout design. We propose \textbf{LayoutDSL}, a novel LLM-based framework for learning an interior layout policy in a domain-specific language (DSL) action space. The DSL provides an explicit symbolic representation of layout information and serves as a structured action space for layout reasoning, where each action corresponds to an interpretable design decision. Under this DSL-based policy learning paradigm, we construct 3D-FrontDSL, a dataset of room-structure annotations paired with synthetic DSL action sequences for supervised fine-tuning. To promote a more generalizable and scalable policy with verifiable feedback, we design rewards grounded in interior design principles and physical plausibility, and optimize the policy via reinforcement learning. Extensive experiments demonstrate that LayoutDSL substantially improves spatial plausibility and design logicality over strong baselines and existing methods.