Back2Struct:让结构化图像重新变得可编辑
Back2Struct: Making Structured Images Editable Again
浏览论文内容
中文总结 AI 辅助
Back2Struct通过从图像中恢复对象级SVG/XML代码,使结构化图像重新可编辑,并利用复合奖励优化,提升准确性、可编辑性和有效性。
中文摘要 AI 辅助
结构化图像,例如图表、流程图和示意图,本质上是符号化的,可以紧凑地表示为可编辑的格式,然而在实践中,它们通常以图像形式呈现,因此无法进行图形化编辑。这种不匹配给研究人员、工程师和设计师带来了重大挑战,他们希望将现有图形内容的修改版本整合到新材料中,而无需手动重建。在本研究中,我们提出了Back2Struct,它通过直接从图像表示中恢复矢量图形代码(SVG/XML)来“让结构化图像重新变得可编辑”。给定一张结构化图形的图像,Back2Struct预测语义级别的对象级SVG/XML代码,该代码显式编码文本、形状、拓扑和布局,而不是执行低级别的像素矢量化。生成的代码可以无缝导入到PowerPoint等工具中,使用户能够在保持结构保真度的同时编辑、细化、重新样式化和重用图形内容。除了在真实SVG令牌序列上进行监督微调外,我们还通过基于奖励的学习进一步优化Back2Struct,以更好地满足部署时的要求:输出应在语法上有效、适当简洁,并在视觉上与输入图表保持一致。具体来说,我们设计了一个复合奖励,共同鼓励SVG/XML的可编译性、与参考代码的长度一致性,以及生成图形与真实图形之间的结构或语义相似性。这些互补信号引导模型生成不仅更接近训练分布,而且在实践中更完整、可编辑和可渲染的SVG。实验表明,Back2Struct在准确性、可编辑性、有效性和用户对齐方面优于基线。数据集和代码可在以下网址获取:this http URL
英文摘要
Structured images, such as diagrams, charts, and flowcharts, are inherently symbolic and can be compactly represented in an editable format, yet in practice, they are often rendered as images, and therefore not graphically editable. This mismatch presents a significant challenge for researchers, engineers, and designers who wish to incorporate modified versions of existing graphic content into new materials without manually reconstructing it. In this study, we presentBack2Struct, which "makes structured images editable again" by directly recovering vector graphics code (SVG / XML) from image representations. Given an image of a structured graphic, Back2Struct predicts semantically object-level SVG / XML code that explicitly encodes text, shapes, topology, and layout, rather than performing low-level pixel vectorization. The generated code can be seamlessly imported into tools such as PowerPoint, allowing users to edit, refine, restyle, and reuse graphic content while preserving structural fidelity. Beyond supervised fine-tuning on ground-truth SVG token sequences, we further optimize Back2Struct with reward-based learning to better match deployment-time requirements: the output should be syntactically valid, properly concise, and visually faithful to the input diagram. Specifically, we design a composite reward that jointly encourages SVG / XML compilability, length consistency with the reference code, and structural or semantic similarity between the generated and ground-truth graphics. These complementary signals guide the model to produce SVGs that are not only closer to the training distribution, but also more complete, editable, and renderable in practice. Experiments show that Back2Struct improves accuracy, editability, validity, and user alignment over baselines. Dataset and code are available at: pengyu965.github.io/Back2Struct.github.io
发表机构
- University at Buffalo, SUNY(纽约州立大学布法罗分校)
机构由 AI 辅助整理,请以论文原文为准。