GlyphAnchor:通过位置锚定字形先验增强视觉文本渲染
GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors
浏览论文内容
中文总结 AI 辅助
针对图像生成与编辑模型的文本渲染难题,提出GlyphAnchor方法,结合字形先验提升文本保真度,引入InfoTextBench基准验证,效果显著。
中文摘要 AI 辅助
对于图像生成与编辑模型而言,渲染准确的文本仍是一项难题,尤其是当目标包含长、复杂且密集排列的文本或生僻字时。现有方法要么通过更强的主干网络和以数据为中心的训练来提升原生文本渲染能力,但未采用显式字形先验;要么通过专门设计融入字形先验,但在挑战性场景下仍不够准确和鲁棒。我们提出GlyphAnchor,一种适用于文本到图像及图像编辑扩散Transformer模型的新型文本渲染增强方法。GlyphAnchor通过轻量字形补丁条件增强主干网络,其位置通过模型原生位置编码锚定到目标图像。我们通过分阶段监督微调训练该能力,并进一步通过文本感知的后训练进行优化以提升鲁棒性。我们还引入了InfoTextBench,一个用于评估生成与编辑场景下文本丰富视觉文本渲染的基准。在多个主干网络及基准上的实验,包括长、复杂、密集排列文本和生僻字场景,显示GlyphAnchor在保持整体图像质量的同时,持续提升文本保真度。
英文摘要
Rendering accurate text remains difficult for image generation and editing models, especially when the target contains long, complex, and densely arranged text or rare characters. Existing approaches either improve native text rendering through stronger backbones and data-centric training without explicit glyph priors, or incorporate glyph priors through specialized designs that remain insufficiently accurate and robust under challenging scenarios. We introduce GlyphAnchor, a novel text-rendering enhancement method for both text-to-image and image-editing diffusion transformer models. GlyphAnchor enhances the backbone with lightweight glyph patch conditions whose positions are anchored to the target image through the model's native positional encoding. We train this capability with staged supervised finetuning and further refine it with text-aware post-training to improve robustness. We also introduce InfoTextBench, a benchmark for evaluating text-rich visual text rendering in both generation and editing settings. Experiments across multiple backbones and benchmarks, including long, complex, and densely arranged text and rare character scenarios, show that GlyphAnchor consistently improves text fidelity while preserving overall image quality.
发表机构
- Fudan University(复旦大学)
- Xiaohongshu Inc.(小红书公司)
- Shanghai Innovation Institute(上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。