基于中间渲染的强化学习用于图像到代码生成
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
浏览论文内容
中文总结 AI 辅助
本文提出IR4RL框架,利用中间渲染进展作为标记级奖励,为图像到代码生成提供密集监督,在SVG和TikZ任务上超越现有方法,实现新的最先进性能。
中文摘要 AI 辅助
强化学习越来越多地被用于对视觉-语言模型进行后训练,以完成图像到代码的生成任务,例如从参考图像生成SVG代码,其通过优化从最终渲染输出计算得到的奖励来实现。然而,依赖单一的最终奖励所提供的反馈是稀疏的,并且与单个标记的贡献对齐不佳。一个生成的程序可能包含能够准确再现目标图像某些部分的操作,同时也有引入错误的其他操作,但所有标记都根据相同的最终结果进行训练。我们观察到,许多中间代码前缀不仅可执行,而且已经能产生有意义的局部渲染,这些渲染反映了向目标进展的情况。这一特性为生成过程提供了更密集监督的自然来源。基于这一观察,我们提出了IR4RL,一个具有标记级渲染进展奖励的强化学习框架,该框架将中间渲染之间的变化转化为针对生成序列的局部反馈。我们在图像到SVG和图像到TikZ生成任务上评估了我们的方法。在这两项任务中,我们的方法都优于监督微调和标准GRPO,产生了新的开源最先进模型。这表明中间渲染为图像到代码模型的强化学习后训练提供了一种简单而有效的过程监督来源。
英文摘要
Reinforcement learning is increasingly used to post-train vision-language models for image-to-code generation, such as generating SVG code from a reference image, by optimizing rewards computed from the final rendered output. However, relying on a single terminal reward provides sparse feedback that is poorly aligned with the contribution of individual tokens. A generated program may contain operations that accurately reproduce some parts of the target image alongside others that introduce errors, yet all tokens are trained from the same final outcome. We observe that many intermediate code prefixes are not only executable, but already produce meaningful partial renders that reflect progress toward the target. This property provides a natural source of denser supervision during generation. Based on this observation, we introduce IR4RL, an RL framework with a token-level render-progress reward that turns changes between intermediate renders into localized feedback for the generated sequence. We evaluate our approach on Image-to-SVG and Image-to-TikZ generation. Across both tasks, our method improves over supervised fine-tuning and standard GRPO, yielding new state-of-the-art open-source models. This shows that intermediate rendering provides a simple and effective source of process supervision for RL post-training of image-to-code models.
发表机构
- Weizmann Institute of Science(魏茨曼科学研究所)
- MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。