arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.11640cs.CVcs.AI

分词使多模态大语言模型能够理解、生成和编辑建筑平面图

Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plans

Sizhong Qin, Ramon Elias Weber, Xinzheng Lu

首次发表 更新
浏览论文内容

中文总结 AI 辅助

HouseMind通过引入离散房间实例标记,实现了建筑平面图的统一理解和生成,具备高效且可控的布局生成能力。

中文摘要 AI 辅助

建筑平面图设计需要在几何、语义和空间层次之间进行联合推理,这仍然是当前AI系统的主要挑战。尽管最近的扩散模型和语言模型提高了视觉保真度,但它们仍然在空间推理和可控生成方面存在困难。我们提出了HouseMind,这是一种多模态大语言模型,它在一个框架中统一了平面图的理解、生成和编辑。我们引入了离散的房间实例标记,以构建一个统一的词汇表,该词汇表连接布局和符号推理。通过多模态对齐和指令微调,该模型能够从文本指令中合成连贯且可控的布局。实验展示了该框架如何在保持高效和本地部署的同时,实现优越的几何有效性和可控性。

英文摘要

Architectural floor plan design demands joint reasoning over geometry, semantics, and spatial hierarchy, which remains a major challenge for current AI systems. Although recent diffusion and language models improve visual fidelity, they still struggle with coherent spatial reasoning and controllable generation. We present HouseMind, a multimodal large language model that unifies floor plan understanding, generation, and editing in one framework. We introduce discrete room-instance tokens to construct a unified vocabulary that bridges layouts and symbolic reasoning. With multimodal alignment and instruction tuning, the model synthesizes coherent, controllable layouts from text instructions. Experiments show how the framework achieves superior geometric validity and controllability while remaining efficient and locally deployable.

发表机构

  • Tsinghua University(清华大学)
  • UC Berkeley(伯克利大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑