arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23386cs.CVcs.GR

ProxyBuild:基于网格锚定程序化代理的文本引导结构化三维建筑生成

ProxyBuild: Text-Guided Structured 3D Building Generation with Mesh-Anchored Procedural Proxies

  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
  • Pengcheng Laboratory(鹏城实验室)
  • Harbin Institute of Technology(哈尔滨工业大学)
  • Harbin Institute of Technology, Suzhou Research Institute(哈尔滨工业大学苏州研究院)

机构由 AI 辅助整理,请以论文原文为准。

Xiang Tang, Ruotong Li, Xiaopeng Fan

AI总结:

提出ProxyBuild混合框架,利用网格锚定程序化代理(MAPP)解耦生成过程,结合面-边二分图编码器和LLM解析,实现文本引导的结构化三维建筑生成,显著缓解过平滑、碰撞等问题。

AI中文摘要:

文本引导的三维建筑生成具有巨大的应用潜力,然而现有的生成模型通常输出不可分离的单一网格或不可交互的渲染表示。虽然程序化建模可以生成具有层次结构的可编辑建筑,但规则编写费时费力,即使借助大语言模型(LLMs),在几何约束下有效求解程序化规则仍然具有挑战性。本文提出ProxyBuild,一种用于结构化建筑生成的新型混合框架。我们引入网格锚定程序化代理(MAPP)作为新型中间表示,将建筑组件紧密锚定到几何外壳上,从而将生成任务解耦为两个阶段:代理预测和代理到资产的实例化。首先,我们构建带有MAPP标注的建筑数据集,用于训练我们设计的面-边二分图编码器。通过在异构网格图上显式建模拓扑元素的特征交互,该编码器能准确推断面和边的语义角色。随后,以LLMs解析的文本风格和属性参数为条件,我们通过集成带有硬约束的空间放置逻辑,实现高精度的资产检索和组装。大量实验表明,ProxyBuild不仅显著缓解了建筑生成中常见的过平滑、组件碰撞和结构损坏等问题,还能准确解析来自不同来源的无语义外壳。在多种指标上优于先前基线,我们的方法能够从文本稳健地生成结构清晰、细节丰富且可后期编辑的三维建筑,从而为虚拟现实和数字孪生等下游应用提供可靠且交互式的内容基础。

英文摘要:

Text-guided 3D building generation holds tremendous application potential, yet existing generative models typically output inseparable single meshes or non-interactive rendered representations. While procedural modeling can generate editable buildings with hierarchical structures, rule authoring is laborious, and even with the aid of large language models (LLMs), it remains challenging to effectively solve procedural rules under geometric constraints. In this paper, we propose ProxyBuild, a novel hybrid framework for structured building generation. We introduce the Mesh-Anchored Procedural Proxy (MAPP) as a novel intermediate representation, which tightly anchors building components onto geometric shells, thereby decoupling the generation task into two phases: proxy prediction and proxy-to-asset instantiation. First, we construct a building dataset with MAPP annotations to train our designed face-edge bigraph encoder. By explicitly modeling the feature interactions of topological elements on heterogeneous mesh graphs, this encoder accurately infers the semantic roles of faces and edges. Subsequently, conditioned on textual styles and attribute parameters parsed by LLMs, we accomplish high-precision asset retrieval and assembly by integrating a spatial placement logic with hard constraints. Extensive experiments show that ProxyBuild not only significantly mitigates common issues in building generation such as over-smoothing, component collisions, and structural corruptions, but also accurately parses semantic-free shells from diverse sources. Outperforming prior baselines across various metrics, our method can robustly generate structurally clear, detail-rich, and post-editable 3D buildings from text, thereby providing a reliable and interactive content foundation for downstream applications such as virtual reality and digital twins.

补充信息

↑