发表机构
Autodesk Research(Autodesk研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FLOORA提出了一种结合领域专用DSL、定制分词、预训练、SFT和RL的小型语言模型,用于建筑布局生成,以0.6B参数在VLM和人工评估中超越大型模型,验证了领域专用配方在结构化输出工程领域的有效性。
AI 中文摘要
基础模型是强大的生成器,但许多工程领域需要通用系统难以处理的结构化表示。我们引入了FLOORA(基于强化学习对齐的楼层布局优化),一个用于建筑布局生成的小型领域专用语言(DSL)模型系列。凭借专门的数据和对齐,我们的0.6B模型优于大得多的前沿模型,在分布外真实世界建筑上实现了高达92.0%的VLM评判胜率,在合成建筑上实现了96.0%的胜率。人工评估进一步证实了这些结果,FLOORA在89.3%的评估中被选为最佳模型。FLOORA结合了令牌高效的DSL、自定义分词、领域专用预训练、监督微调(SFT)以及带有学习到的人类偏好和可验证奖励的强化学习(RL)。这一流程提高了建筑和几何有效性,并得到了广泛的实证评估和消融研究的支持。尽管专注于建筑领域,我们的结果表明,类似的领域专用配方可能在其他具有结构化、可验证输出的工程领域中发挥作用。数据集、模型和推理代码可在该https URL获取。
英文摘要
Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly. We introduce FLOORA (Floor Layout Optimization with RL Alignment), a family of small domain-specific language (DSL) models for architectural layout generation. With specialized data and alignment, our 0.6B model outperforms much larger frontier models, achieving VLM judge win rates up to 92.0% on out-of-distribution real-world buildings and 96.0% on synthetic buildings. Human evaluations further corroborate these results, with FLOORA selected as the best model in 89.3% of evaluations. FLOORA combines a token-efficient DSL, custom tokenization, domain-specific pretraining, supervised fine-tuning (SFT), and reinforcement learning (RL) with learned human-preference and verifiable rewards. This pipeline improves architectural and geometric validity, supported by extensive empirical evaluation and ablation studies. Although focused on architecture, our results suggest that similar domain-specific recipes may be useful in other engineering domains with structured, verifiable outputs. Datasets, models, and inference code are available at https://github.com/AutodeskAILab/floora.