arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19000cs.CV

场景布置(Mise-en-Scène):用于人机协同设计共创的扩散Transformer中隐式布局的涌现

Mise-en-Scène: Implicit Layout Emergence in Diffusion Transformers for Human-AI Design Co-Creation

Zipeng Xu, Ryan Murdock, Umberto Michieli

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Mise-en-Scène框架,通过微调扩散Transformer实现隐式布局涌现,结合匹配放置步骤保证素材保真度,在PrismLayersPlus基准上生成的设计感知质量显著优于现有方法。

中文摘要 AI 辅助

从用户提供的元素自动生成平面设计,既需要整体构图的连贯性,又要精确保留每个素材。现有方法通过语言模型预测显式边界框坐标作为布局,再将素材粘贴到对应位置,这种方式将空间规划与视觉合成分离,易产生僵硬、比例失调的构图。我们转而探究布局是否能在预训练的图像编辑扩散Transformer中隐式涌现。本文提出Mise-en-Scène,这是一个两阶段框架:第一阶段,通过少量精选LoRA(低秩适配)微调的扩散Transformer生成完整设计草稿,其中元素的排列与渲染画布共同涌现;第二阶段,采用确定性匹配放置步骤,将原始高分辨率素材移动到草稿位置,保证素材的精确保真度,生成可编辑的分层设计,供设计师进一步优化,而非仅输出平面图像。值得注意的是,仅需对预训练Transformer进行最小程度的适配即可满足需求,无需为多元素生成引入常用的专用条件机制。在大规模PrismLayersPlus基准测试中,Mise-en-Scène生成的设计在感知质量上与真实值最为接近,大幅优于LLM布局规划器和专用布局Transformer,而我们的匹配放置阶段则缩小了与真实合成图的剩余保真度差距。

英文摘要

Automating graphic design synthesis from user-provided elements requires both a coherent overall composition and the exact preservation of each asset. Existing methods predict a layout as explicit bounding-box coordinates with a language model and then paste the assets into it, which separates spatial planning from visual synthesis and tends to produce rigid, mis-scaled compositions. We instead ask whether the layout can emerge implicitly inside a pretrained image-editing diffusion transformer. We present Mise-en-Scène, a two-stage framework. In the first stage, a diffusion transformer adapted with a small, knockout-selected LoRA drafts a complete design in which the arrangement of the elements emerges jointly with the rendered canvas. In the second stage, a deterministic match-and-place step moves the original high-resolution assets to the drafted positions, which guarantees exact asset fidelity and yields an editable, layered design that a designer can keep refining rather than a flat image. Notably, a minimal adaptation of the pretrained transformer already suffices, without the specialized conditioning machinery commonly introduced for multi-element generation. On the large-scale PrismLayersPlus benchmark, the designs produced by Mise-en-Scène are the closest to the ground truth in perceived quality among all compared methods, by a wide margin over both an LLM layout planner and a specialized layout transformer, while our match-and-place stage bridges the remaining fidelity gap to the ground-truth composites.

发表机构

  • Canva Research(Canva研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑