发表机构
University of Massachusetts Amherst; University of Central Florida(马萨诸塞大学阿默斯特分校; 中佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估结合视觉反馈、结构化编辑与约束解码的约束迭代优化方法,探索用现成VLMs生成SVG的能力,发现约束解码提升编译成功率但VLMs存在视觉推理不足的局限。
AI 中文摘要
可缩放矢量图形(SVGs)支撑着现代视觉生态系统的大部分内容,但当前最先进的生成模型几乎完全聚焦于光栅图像。本研究探索推理时方法能否解锁现成视觉语言模型(VLMs)的SVG生成能力。我们系统评估了一种结合视觉反馈、结构化编辑与约束解码的约束迭代优化框架,以明确当前VLMs在SVG生成方面的能力与局限。在多个VLMs及不同生成设置下,我们发现约束解码可提升编译成功率,而迭代优化则暴露出视觉推理与自我修正能力的不足。研究结果凸显了使用推理时方法适配通用VLMs进行SVG生成的潜力与当前局限。
英文摘要
Scalable Vector Graphics (SVGs) power much of the modern visual ecosystem, yet state-of-the-art generative models focus almost entirely on rasterized images. We explore whether inference-time methods can unlock SVG generation capabilities in off-the-shelf vision-language models (VLMs). We systematically evaluate a constrained iterative refinement harness that combines visual feedback, structured editing, and constrained decoding to characterize the capabilities and limitations of current VLMs for SVG generation. Across multiple VLMs and generation settings, we find that constrained decoding improves compilation success rates, while iterative refinement reveals a deficit in visual reasoning and self-correction. Our results highlight both the promise and current limitations of using inference-time methods to adapt general-purpose VLMs for SVG generation.
CommentsAccepted to Pacific Graphics 2026 Poster Track. 2 Pages. 2 Figures. 1 Table