AI 中文总结
本文提出PosterText框架,将文本块作为原子单元实现电商海报的统一生成与编辑,构建了带块级标注的数据集与评估基准,经实验验证其性能优于现有相关方法。
AI 中文摘要
自动化电商海报设计既需要高质量的海报生成,也需要对现有设计进行灵活编辑。然而,现有大多数方法要么针对端到端的海报生成,要么遵循多阶段设计流程,对现有海报的灵活且精确编辑能力有限。为实现电商海报的统一生成与编辑,本文提出文本块生成与编辑,这是一种将文本块视为原子单元的统一任务表述,涵盖海报生成、块添加、块删除、块修改四种操作,支持可选的参考引导风格控制。基于此,本文提出PosterText模型,该模型采用四阶段课程训练,包括文本渲染预训练、指令跟随训练、用于偏好对齐的强化学习,以及用于执行优化的空间引导自蒸馏。本文还构建了一个具有块级标注的大规模数据集和一套综合评估基准。大量实验表明,PosterText在与现有生成和编辑方法的对比中取得了具有竞争力的性能,验证了所提框架的有效性。
英文摘要
Automated e-commerce poster design requires both high-quality poster generation and flexible editing of existing designs. However, most existing methods either target end-to-end poster generation or follow multi-stage design pipelines, with limited capability for flexible and precise editing of existing posters. To enable unified generation and editing of e-commerce posters, we introduce Text Patch Generation and Editing, a unified task formulation that treats text patches as atomic units and covers four operations: poster generation, patch addition, patch deletion, and patch modification, with optional reference-guided style control. Based on this, we propose PosterText, a unified model trained with a four-stage curriculum, including text rendering pretraining, instruction-following training, reinforcement learning for preference alignment, and spatial guidance self-distillation for execution refinement. We further construct a large-scale dataset with patch-level annotations and a comprehensive benchmark for evaluation. Extensive experiments demonstrate that PosterText achieves competitive performance against existing generation and editing approaches, validating the effectiveness of the proposed framework.