arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RefineSVG:视觉反馈驱动的图像转SVG生成强化学习

RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation

Shaobo Liu, Feiqiao Mao, Shuaishuai Zhou, Yan Zhan, Weiqi Tan, Zhiqiong Lu, Zhengping Liang

arXiv 2607.27699首次发表:更新:

发表机构

Shenzhen University; Peking University(深圳大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RefineSVG是一种视觉反馈驱动的闭环框架,通过引入SVG语义词汇表和渐进式训练,使MLLM实现高保真图像转SVG生成,性能优于现有基线。

AI 中文摘要

我们提出RefineSVG,这是一种单步闭环视觉反馈框架,可使多模态大语言模型(MLLM)通过自校正实现高保真的图像转SVG生成。现有基于MLLM的方法依赖单次开环推理,模型仅接收一次视觉输入,就必须生成数千个SVG代码标记且无中间验证,这种范式不可避免地会在复杂图像上产生几何漂移、误差累积和视觉幻觉。RefineSVG通过在初始SVG生成阶段后调用外部渲染引擎,将渲染输出与目标图像进行对比,从而克服这一局限;对比生成的多维视觉残差图(Diff-Map)会以ReAct风格的校正信号形式反馈给模型,驱动针对性的校正步骤。为支持这种渲染-观察-校正的交互,我们还引入了面向SVG的语义词汇表,可将标记序列压缩超过52%。涵盖监督微调、拒绝采样冷启动数据构建和端到端智能体强化学习的渐进式训练流程,使模型与闭环视觉校正对齐。大量实验表明,RefineSVG在重建保真度、结构准确性和代码质量上始终优于现有基线。代码可在指定链接获取。

英文摘要

We propose RefineSVG, a single-step closed-loop visual feedback framework that enables multimodal large language models (MLLMs) to perform high-fidelity image-to-SVG generation through self-correction. Existing MLLM-based approaches rely on single-pass open-loop inference, where the model receives visual input only once and must generate thousands of SVG code tokens without intermediate verification. This paradigm inevitably leads to geometric drift, error accumulation, and visual hallucination on complex images. RefineSVG overcomes this limitation by invoking an external rendering engine after an initial SVG generation pass to compare the rendered output against the target image. The comparison yields a multi-dimensional visual residual map (Diff-Map) that is fed back to the model as a ReAct-style correction signal, driving a targeted correction step. To support this render-observe-correct interaction, we further introduce an SVG-oriented semantic vocabulary that compresses token sequences by over 52%. A progressive training pipeline spanning supervised fine-tuning, rejection-sampling cold-start data construction, and end-to-end agentic reinforcement learning aligns the model with closed-loop visual correction. Extensive experiments show that RefineSVG consistently outperforms existing baselines in reconstruction fidelity, structural accuracy, and code efficiency.Code is available at https://github.com/liuxiaobo66/RefineSVG.

Comments17 pages, 5 main-paper figures. Accepted at the 34th ACM International Conference on Multimedia (ACM MM 2026). Includes the complete supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑