arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态思维与可渲染程序

Multimodal Thinking with Renderable Programs

Sunli Chen, Ding Zhong, Ziqiao Ma, Jiaxin Liu, Zeyuan Yang, Hao Zhang, Lie Lu, Joyce Chai, Chuang Gan

arXiv 2609.30130首次发表:更新:

发表机构

University of Massachusetts Amherst; University of Michigan; University of Illinois Urbana-Champaign; Dolby Laboratories(马萨诸塞大学阿默斯特分校; 密歇根大学; 伊利诺伊大学厄巴纳-香槟分校; 杜比实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SVGLM利用SVG原语连接文本与图像,使通用视觉语言模型能在推理中生成图像,并在数学推理基准上展现强大性能。

AI 中文摘要

当前的视觉语言模型(VLMs)在视觉内容理解和基于文本的推理方面表现出色,但其结构限制了将图像融入推理链的进展。尽管全模态模型在统一文本和图像生成方面做出了努力,但它们专注于开放领域的视觉任务,由于图像的光栅化或潜在表示而缺乏可操作性。我们引入了SVGLM,一个利用可缩放矢量图形(SVG)原语在推理任务中连接文本和图像的框架。我们利用SVG作为图像描述和文本指令的双重性,提供了一种更紧凑、可解释的解决方案,使通用VLMs具备在推理过程中生成图像的能力。我们提供了一个大型精选的基于SVG的图像编辑数据集,以及调整开源VLMs的范式。在数学推理基准上的实验表明,SVGLM实现了强大的SVG生成能力以及思维伴随图像的智能。我们的结果突出了SVG作为构建更健壮的数字领域代理的合适媒介,弥合了基于文本的思维与基于像素的图像之间的差距。

英文摘要

Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorporating images into the reasoning chain. Though Omnimodal models have made efforts in unifying text and image generation, they focus on visual tasks in the open-domain, lacking tractability due to rasterized or latent representations of images. We introduce SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks. We exploit the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generating images within the reasoning process. We provide a large curated dataset of SVG-based image editing dataset, as well as the paradigm to tune open-source VLMs. Experiments on a mathematical reasoning benchmark demonstrate that SVGLM achieves strong SVG generation power as well as think-with-image intelligence. Our results highlight SVG as a suitable medium for building more robust digital domain agents, bridging the gap between text-based thinking and pixel-based images.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑