arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2311.16974cs.CV

COLE:面向多层可编辑图形设计的分层生成框架

COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design

Peidong Jia, Chenxuan Li, Yuhui Yuan, Zeyu Liu, Yichao Shen, Bohan Chen, Xingru Chen, Yinglin Zheng, Dong Chen, Ji Li, Xiaodong Xie, Shanghang Zhang, Baining Guo

首次发表 更新
浏览论文内容

中文总结 AI 辅助

COLE系统通过分层任务分解,将模糊意图转化为高质量多层图形设计,集成微调LLM、LMM和扩散模型,并构建基准验证其优越性,支持灵活编辑。

中文摘要 AI 辅助

图形设计自15世纪以来不断发展,在广告中扮演着至关重要的角色。高质量设计的创作需要面向设计的规划、推理和逐层生成。与近期将GPT-4与现有设计模板集成以构建定制GPT的CanvaGPT不同,本文介绍了COLE系统——一个旨在全面应对这些挑战的分层生成框架。该COLE系统能够将模糊的意图提示转换为高质量的多层图形设计,同时支持基于用户输入的灵活编辑。此类输入的示例可能包括诸如“为久石让的音乐会设计一张海报”之类的指令。关键洞察在于将文本到设计生成这一复杂任务分解为一系列更简单的子任务层级,每个子任务由专门模型协作处理。这些模型的结果随后被整合以产生连贯的最终输出。我们的分层任务分解能够简化复杂流程并显著提升生成可靠性。我们的COLE系统包含多个微调的大语言模型(LLMs)、大型多模态模型(LMMs)和扩散模型(DMs),每个模型均专门针对设计感知的逐层字幕生成、布局规划、推理以及图像和文本生成任务进行定制。此外,我们构建了DESIGNINTENTION基准,以证明我们的COLE系统在从用户意图生成高质量图形设计方面优于现有方法。最后,我们呈现了一个类似Canva的多层图像编辑工具,以支持对生成的多层图形设计图像的灵活编辑。我们将COLE系统视为迈向未来解决更复杂和多层图形设计生成任务的重要一步。

英文摘要

Graphic design, which has been evolving since the 15th century, plays a crucial role in advertising. The creation of high-quality designs demands design-oriented planning, reasoning, and layer-wise generation. Unlike the recent CanvaGPT, which integrates GPT-4 with existing design templates to build a custom GPT, this paper introduces the COLE system - a hierarchical generation framework designed to comprehensively address these challenges. This COLE system can transform a vague intention prompt into a high-quality multi-layered graphic design, while also supporting flexible editing based on user input. Examples of such input might include directives like ``design a poster for Hisaishi's concert.'' The key insight is to dissect the complex task of text-to-design generation into a hierarchy of simpler sub-tasks, each addressed by specialized models working collaboratively. The results from these models are then consolidated to produce a cohesive final output. Our hierarchical task decomposition can streamline the complex process and significantly enhance generation reliability. Our COLE system comprises multiple fine-tuned Large Language Models (LLMs), Large Multimodal Models (LMMs), and Diffusion Models (DMs), each specifically tailored for design-aware layer-wise captioning, layout planning, reasoning, and the task of generating images and text. Furthermore, we construct the DESIGNINTENTION benchmark to demonstrate the superiority of our COLE system over existing methods in generating high-quality graphic designs from user intent. Last, we present a Canva-like multi-layered image editing tool to support flexible editing of the generated multi-layered graphic design images. We perceive our COLE system as an important step towards addressing more complex and multi-layered graphic design generation tasks in the future.

发表机构

  • Microsoft Research Asia(微软亚洲研究院)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑