发表机构
ShanghaiTech University(上海科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对设计模板生成中背景与布局依赖缺失的问题,提出InterIL模型,通过可学习通信模块连接冻结的预训练图像与布局扩散模型,联合生成图像和布局,并支持测试时用户偏好引导,显著提升生成和谐度。
AI 中文摘要
本文研究了图形设计模板创建问题,该任务旨在根据输入文本生成背景图像以及背景上前景元素的布局,以形成和谐的构图。以往关于图形设计生成的工作大多采用顺序范式,即按顺序生成设计元素。我们认为,这种顺序方案无法忠实捕捉背景与布局之间的依赖关系(从而无法捕捉联合图像-布局分布),这限制了生成设计模板的质量。为克服这一局限,我们提出了一种名为InterIL的模型,该模型在单一生成过程中联合生成背景图像和布局这两种模态。我们联合模型的新颖设计在于,通过一个可学习的通信模块连接预训练的图像扩散模型和布局扩散模型的骨干网络,以显式建模双向的图像-布局交互。在训练期间,图像和布局骨干网络保持冻结,以维持并利用大量的预训练单模态先验知识,仅更新通信模块,从而使模型能够专注于学习图像-布局交互,从而更好地捕捉联合图像-布局分布,以改善构图和谐性。我们的模型没有特定于设计的归纳偏置,这使其能够更好地保留真实设计的原始特征。我们进一步引入了一种测试时引导策略,使用户能够将其特定偏好施加于生成结果上。我们的实验表明,与先前方法相比,我们的模型在图像、布局以及图像-布局协调性方面能够生成显著更好的结果,产生更接近真实样本的输出。我们还展示了模型在推理时无需重新训练即可强制执行用户偏好的灵活性。
英文摘要
In this paper, we address the problem of graphic design template creation, which generates a background image and a layout of foreground elements over the background to form a harmonious composition from an input text. Prior work on graphic design generation mostly adopts a sequential paradigm, where design elements are generated sequentially. We argue that such a sequential scheme falls short of faithfully capturing the dependency between the background and layout (and thus the joint image-layout distribution), which limits the quality of generated design templates. To overcome this limitation, we propose a model, InterIL, which jointly generates the two modalities, background image and layout, in a single generative process. The novel design of our joint model connects the backbones of pretrained image and layout diffusion models with a learnable communication module to explicitly model bidirectional image-layout interaction. During training, the image and layout backbones are frozen to maintain and leverage the vast pretrained single-modality prior knowledge, while only the communication module is updated, so that the model can focus on learning image-layout interaction and thereby better capture the joint image-layout distribution for improved composition harmony. Our model has no design-specific inductive bias, which allows it to better preserve the original characteristics of realistic designs. We further introduce a test-time guidance strategy to enable users to impose their specific preferences on generated results. Our experiments show that, compared with prior approaches, our model can generate significantly better results in terms of image, layout and image-layout harmonization, producing outputs closer to real samples. We also demonstrate the flexibility of our model in enforcing user preferences at inference without retraining.
CommentsMain paper with supplementary material. Submitted to IEEE Transactions on Visualization and Computer Graphics