发表机构
Institute of Trustworthy Embodied AI, Fudan University; Shanghai Key Laboratory of Multimodal Embodied AI(复旦大学可信具身智能研究院; 上海市多模态具身智能重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对中文文本渲染中OCR奖励忽略字形组合结构的问题,提出IDSpect方法,利用IDS分解与细粒度奖励,在不改生成器且无额外推理成本下,提升结构质量与语义对齐。
AI 中文摘要
对于文本到图像模型而言,渲染精确的中文文本仍然具有挑战性。现有的基于OCR的强化学习奖励机制将解码出的文本与目标字符串进行比较。这类奖励忽视了中文书写的组合性质:一个表意文字由可复用的组件通过明确的空间关系排列而成,然而OCR却将其作为一个原子字符进行评估。因此,视觉上不同的部首级错误可能会得到同样粗略的反馈,这鼓励了仅仅在字形上近似目标而非忠实再现其内部结构的生成结果。我们采用表意文字描述序列(IDS),该序列包含空间运算符和字符组件,并训练一个专家IDS识别器,将渲染出的中文文本转录为这种表示。基于该识别器,我们引入了IDSpect,它确定性地将目标文本分解为IDS标记,并将裁剪级别的视觉IDS预测与目标序列对齐。全局唯一的标记信用使得这种比较对检测到的文本区域的顺序具有鲁棒性。结合整字符语义奖励,IDSpect提供了具有组件和空间关系的细粒度信用,而无需改变图像生成器或增加推理时间成本。使用GRPO对Qwen-Image进行后训练的实验表明,IDSpect在LongText和GenTextEval上达到了领先的结构质量和语义对齐。
英文摘要
Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.