arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TextRefine:提升产品海报文本编辑中的文本保真度、空间位置与字形渲染

TextRefine: Improving Textual Fidelity, Spatial Placement, and Glyph Rendering for Text Editing in Product Posters

Honglie Wang, Jia Sun, Zijun Li, Junlong Wu, Pengcheng Wei, Jiyuan Wang, Yongrui Heng, Boheng Zhang, Huaiqing Wang, Dewen Fan, Qianqian Gan, Fan Yang, Tingting Gao, Yan-Ming Zhang

arXiv 2608.19637首次发表:更新:

AI 中文总结

研究针对产品海报文本编辑中通用模型的缺陷,提出TextRefine框架与OpenTextEdit数据集,在文本保真度、位置可靠性等指标上优于基线方法,提升了产品海报文本编辑效果。

AI 中文摘要

产品海报中的文本编辑需插入新文本或替换现有文本,同时保留产品外观、背景内容与全局构图。尽管基于指令的图像编辑近期取得进展,但通用模型在该场景下仍不可靠:常遗漏或错误渲染目标文本,将其置于突出产品或已有内容之上,且生成结构失真或视觉不一致的字形。我们提出TextRefine,一种任务对齐的后训练框架,结合监督微调与操作特定的奖励优化以解决这些互补的失败模式。对于文本插入,我们的文本跨度级奖励联合评估语义保真度与目标跨度覆盖率,惩罚与产品及现有文本的空间冲突,并采用门控结构约束保留非文本区域。对于文本替换,我们的字形级奖励利用目标字符的连接时序分类(CTC)后验,为细粒度缺陷提供分级监督,包括笔画缺失、结构变形及视觉相似字符混淆。我们进一步提出OpenTextEdit,一个包含10万张产品海报文本编辑图像的数据集,具备多文本布局、详细文本属性、产品掩码及具有挑战性的低频字符。对插入与替换的大量实验表明,TextRefine在文本保真度、位置可靠性、字形质量上始终优于所评估的图像编辑基线,同时更好地保留源图像内容。

英文摘要

Text editing in product posters entails inserting new text or replacing existing text while preserving product appearance, background content, and global composition. Despite recent progress in instruction-based image editing, general-purpose models remain unreliable in this setting: they often omit or incorrectly render the target text, place it over salient products or pre-existing content, and produce structurally distorted or visually inconsistent glyphs. We introduce \textbf{TextRefine}, a task-aligned post-training framework that combines supervised fine-tuning with operation-specific reward optimization to address these complementary failure modes. For text insertion, our text-span-level reward jointly assesses semantic fidelity and target-span coverage, penalizes spatial conflicts with products and existing text, and employs a gated structural constraint to preserve non-text regions. For text replacement, our glyph-level reward leverages the connectionist temporal classification (CTC) posterior of the target character to provide graded supervision for fine-grained defects, including missing strokes, structural deformations, and confusion among visually similar characters. We further introduce \textbf{OpenTextEdit}, a dataset comprising 100K images for text editing in product posters, with multi-text layouts, detailed text attributes, product masks, and challenging low-frequency characters. Extensive experiments on both insertion and replacement demonstrate that TextRefine consistently outperforms the evaluated image editing baselines in textual fidelity, placement reliability, and glyph quality while better preserving source-image content.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑