arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SG-Layout:基于结构化场景图引导的大语言模型布局生成方法

SG-Layout: Structured Scene Graph-Guided Layout Generation with LLMs

Junsheng Wang, Chao Chen, Mengying Xie, Mingyan Li, Fuqiang Gu

arXiv 2608.01106首次发表:更新:

发表机构

Chongqing University(重庆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SG-Layout通过两阶段训练将结构化空间知识融入LLMs,在三类任务中提升了空间推理与几何一致性,尤其适配关系密集的复杂场景。

AI 中文摘要

从自然语言理解和生成空间连贯的布局,对大语言模型(LLMs)而言是一项基础却极具挑战性的任务。现有LLMs常难以捕捉物体间明确的几何关系与结构依赖。为解决该问题,我们提出SG-Layout,一种图引导的布局生成框架,将结构化空间知识明确融入LLMs。SG-Layout遵循两阶段训练范式:(1)图-语言特征对齐阶段,训练关系图编码器与投影器,将场景图嵌入映射至LLMs的语言空间;(2)指令微调阶段,基于LoRA的适配器实现指令驱动布局生成的高效微调,同时保留主干模型参数固定。我们在图像布局生成、室内场景合成及机器人物体重排任务中评估SG-Layout,实验结果显示,相较于紧凑的开源主干模型,SG-Layout提升了空间推理准确率与几何一致性,在关系密集、结构复杂的场景中优势尤为显著,这些结果凸显了图结构化特征对齐对增强可控布局生成的有效性。

英文摘要

Understanding and generating spatially coherent layouts from natural language remains a fundamental yet challenging task for large language models (LLMs). Existing LLMs often struggle to capture explicit geometric relationships and structural dependencies between objects. To address this issue, we propose SG-Layout, a graph-guided layout generation framework that explicitly incorporates structured spatial knowledge into LLMs. SG-Layout follows a two-stage training paradigm: (1) a graph-language feature alignment stage, where a relational graph encoder and a projector are trained to map scene-graph embeddings into the LLM's linguistic space; and (2) an instruction tuning stage, where LoRA-based adapters enable efficient fine-tuning for instruction-driven layout generation while keeping the backbone frozen. We evaluate SG-Layout on image layout generation, indoor scene synthesis and robotic object rearrangement tasks. Experimental results show that SG-Layout improves spatial reasoning accuracy and geometric consistency over the compact open-source backbone, with particularly clear advantages in relation-dense and compositionally complex scenes. These results highlight the effectiveness of graph-structured feature alignment for enhancing controllable layout generation.

Comments16 pages, 5 figures. Accepted at WAICA 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑