从提示到合成:用于海报生成的空间画布界面
From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
浏览论文内容
中文总结 AI 辅助
该研究提出空间画布界面及适配预训练图像编辑模型的Compo海报生成模型,支持直接推理与智能体模式,通过自动构建监督数据训练,在保持高视觉质量的同时提升了海报生成的构图可控性。
中文摘要 AI 辅助
文本提示是海报生成的间接界面,要求用户将本质上为二维的构图意图编码为一维的单词序列。我们引入一种空间画布界面,使用户能够通过四种互补的绑定类型(语义、身份、文本和像素)在空间中直接合成生成意图,同时为各个元素和全局外观提供文本规范。基于该界面,我们开发了Compo,这是一种从预训练图像编辑模型适配而来的海报生成模型,用于理解空间画布输入和文本规范。Compo支持两种模式:直接推理模式,用户显式构建画布;智能体模式,将高级请求自动转换为规划好的空间画布。为训练Compo,我们开发了可扩展的流水线,自动构建不同绑定类型及其组合的监督数据,实现高效适配,无需从头训练专用海报生成器。我们还引入了一个基准,用于评估对单个绑定类型及其联合构图的遵循情况。实验表明,Compo在保持高视觉质量的同时,比通用图像生成模型和专用海报生成系统实现了更强的构图可控性。通过将意图规范与视觉生成解耦,我们的工作将海报生成从提示转向合成。
英文摘要
Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words. We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance. Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications. Compo supports both direct inference, where users explicitly construct the canvas, and agentic mode, where a high-level request is automatically translated into a planned Spatial Canvas. To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch. We further introduce a benchmark that evaluates adherence to individual binding types and their joint composition. Experiments show that Compo achieves stronger compositional controllability than both general-purpose image generation models and dedicated poster generation systems while maintaining high visual quality. By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.
发表机构
- Fudan University(复旦大学)
- Microsoft Research(微软研究院)
- The University of Sydney(悉尼大学)
- USTC(中国科学技术大学)
- University of Waterloo(滑铁卢大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。