arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22101cs.CV

InnoText:一种用于视觉文本生成和编辑的统一模型

InnoText: A Unified Model for Visual Text Generation and Editing

  • JD.com, Inc.(京东公司)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Chongqing University of Post and Telecommunications(重庆邮电大学)
  • Beijing University of Technology(北京工业大学)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • City University of Hong Kong(香港城市大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Haowei Liu, Runze He, Jian Lu, Ao Ma, Run Ling, Ke Cao, Jiasong Feng, Wei Feng, Shuo Lu, Yexing Xu, Yun Wang, Jing Wang, Zhanjie Zhang

AI总结:

针对视觉文本生成和编辑的挑战,提出基于DiT的统一框架InnoText,通过引入字体大小感知调制模块等策略,构建双语数据集,实现了在单个模型中进行文本生成和编辑,提升了生成精度和编辑质量。

AI中文摘要:

扩散模型在高保真图像合成中取得显著成功,但其在视觉文本生成和编辑方面的应用尚待探索。视觉文本任务对结构规律性和易读性要求高,现有基于UNet的模型生成清晰连贯文本有困难,基于DiT的模型通常限于单一任务。为此提出InnoText,一个基于DiT的统一框架,能在单个模型中执行文本生成和编辑。引入字体大小感知调制模块、小字符感知增强策略和特定任务区域加权损失。构建高质量双语视觉文本数据集。实验结果表明该方法具有卓越的生成精度和编辑质量。

英文摘要:

Diffusion models have recently achieved remarkable success in high-fidelity image synthesis, yet their application to visual text generation and editing remains relatively underexplored. Unlike general image generation, visual text tasks demand precise structural regularity and legibility, which may pose additional challenges for small-scale text and non-Latin scripts such as Chinese. Existing UNet-based models often struggle to produce clear and coherent text, while DiT-based models, though more expressive, are typically limited to a single task, which may lead to redundant training pipelines, inconsistent visual styles, and reduced cross-task generalization. To address these challenges, we propose InnoText, a unified DiT-based framework capable of performing both text generation and editing within a single model. We introduce a Font Size-Aware Modulation (FSAM) module to enhance representations across font scales, a Small-Character Aware Augmentation strategy to improve fine-grained fidelity, and a Task-Specific Region Weighted Loss for adaptive optimization. To support training and evaluation, we also construct a high-quality bilingual (English-Chinese) visual text dataset covering diverse fonts, sizes, and backgrounds. Experimental results demonstrate that our method achieves superior generation accuracy and editing quality, producing visually appealing and realistic text images.

补充信息

相关深度报道

↑