arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过VLM驱动的层次化语义解析进行组合式SVG生成

Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing

Sehwan Park, Taehoon Kim, Geonhee Han, Dohyun Kim, Seung Wook Kim, Paul Hongsuck Seo

arXiv 2609.14657首次发表:更新:

发表机构

Korea University; KAIST AI(高丽大学; 韩国科学技术院人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VLM驱动的智能体框架,通过层次化语义解析递归分解视觉场景,生成结构完整、可编辑的SVG,并引入语义基准和指标,超越现有平面生成方法的分组质量与可编辑性。

AI 中文摘要

尽管视觉语言模型(VLM)在视觉推理方面表现出色,但生成结构化、可编辑的可缩放矢量图形(SVG)仍是一项基本挑战。现有流程主要生成扁平的、语义无关的路径集合,编辑单个对象需要手动识别其组成路径。为解决此问题,我们提出了一种VLM驱动的智能体框架,用于语义组合式SVG生成。我们的流程通过自顶向下的分解、视觉定位和提示驱动的无模态遮挡恢复,递归地将视觉场景解析为语义和几何层次结构,确保每个组件在几何上完整。此外,我们引入了带有手工标注语义组和新型子组件指标(语义召回率/精确率、PERE)的语义SVG基准,以明确评估结构组合性和功能可编辑性。实验表明,我们原生预测的结构在分组质量和可编辑性方面均超越了现有平面生成方法的上限,同时保持了最先进的视觉保真度。

英文摘要

While Vision-Language Models (VLMs) excel at visual reasoning, generating structured, editable Scalable Vector Graphics (SVG) remains a fundamental challenge. Existing pipelines predominantly yield flat, semantically agnostic collections of paths, where editing a single object requires manually identifying its constituent paths. To address this, we propose a VLM-driven agentic framework for semantic compositional SVG generation. Our pipeline recursively parses visual scenes into semantic and geometric hierarchies via top-down decomposition, visual grounding, and prompt-driven amodal occlusion recovery, ensuring each component is geometrically complete. Furthermore, we introduce the Semantic SVG Benchmark with human-annotated semantic groups and novel sub-component metrics (Semantic Recall/Precision, PERE) to explicitly evaluate structural compositionality and functional editability. Experiments show that our natively predicted structures surpass the upper bounds of existing flat-generation methods in both grouping quality and editability, while maintaining state-of-the-art visual fidelity.

Comments26 pages, Accepted to EMNLP 2026 (Main)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑