STEPS:基于扩散模型与对比风格编码的保留风格场景文本编辑
STEPS: Scene Text Editing with Preserved Style Using Diffusion and Contrastive Style Encoding
- Roblox
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
STEPS提出一种扩散模型架构,通过独立于文本内容的风格编码器与多语义条件结合,在场景文本编辑中实现更优的风格保留、可读性和主观质量。
AI中文摘要:
我们提出了保留风格的场景文本编辑(STEPS),一种用于图像中高质量文本替换的新型扩散模型架构。场景文本编辑(STE),也称为视觉文本编辑,包括在保留原始风格(如字体、颜色、方向、背景等)的同时更改图像中的文本内容。STEPS通过专注于改进风格保留,推进了STE领域的最先进水平。我们引入了一个用于视觉文本的风格编码器,该编码器独立于文本内容捕获风格,以及一个将风格编码器与多个语义条件(目标文本字符编码和渲染字形)相结合的模型架构。STEPS在风格保留、输出可读性和主观质量方面均优于先前的STE方法。
英文摘要:
We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation, background, etc. STEPS advances the state of the art in STE through directed focus on improved style preservation. We introduce a style encoder for visual text that captures style independently of textual content, and a model architecture that combines the style encoder with multiple semantic conditions (target text characters encoding and rendered glyphs). STEPS achieves superior results to previous STE methods in style preservation, output readability, and subjective quality.