发表机构
University of Modena and Reggio Emilia(摩德纳和雷焦艾米利亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SLS(SVG隐空间),一种基于Transformer的自编码器,学习SVG路径的稠密可逆隐表示,可降低下游任务FLOPs超150倍,为矢量图形研究提供通用隐空间基础。
AI 中文摘要
可缩放矢量图形(SVG)是与分辨率无关的视觉内容的基础媒介,但深度学习领域一直缺乏针对矢量表示的连续、稠密且可逆的隐空间,而这类隐空间是变分自编码器及其衍生模型长期以来为光栅图像提供的基础构建模块。本文提出SLS(SVG Latent Space,SVG隐空间),一种基于Transformer的自编码器,用于学习单个SVG路径的紧凑稠密表示,SVG路径是可组成任意SVG图像的原子视觉元素。SLS在统一的基于BPE的令牌词汇表中对SVG命令、坐标数据和视觉属性进行建模,学习固定大小的隐表示,该表示同时捕获结构与外观,可解码回具有高保真度的有效、风格一致的SVG路径。所得嵌入空间具有鲁棒性、可逆性和结构性:嵌入位于单位超球面上,可通过简单的向量空间操作实现高效的相似性搜索、组合和下游条件设置。最后,本文证明SLS可在不同任务间泛化,与基于令牌的方法相比,其浮点运算量(FLOPs)降低了150倍以上,为矢量图形研究建立了通用的隐空间基础。
英文摘要
Scalable Vector Graphics are a fundamental medium for resolution-independent visual content, yet the deep learning community lacks a continuous, dense, and invertible latent space for vector representations, the kind of foundational building block that Variational Autoencoders and their descendants have long provided for raster images. We introduce SLS (SVG Latent Space), a Transformer-based autoencoder that learns compact dense representations of individual SVG paths, the atomic visual elements from which any SVG image can be composed. By modeling SVG commands, coordinate data, and visual properties within a unified BPE-based token vocabulary, SLS learns fixed-size latent representations that jointly capture structure and appearance, and can be decoded back into valid, style-consistent SVG paths with high fidelity. The resulting embedding space is robust, invertible, and structured: embeddings lie on a unit hypersphere, enabling efficient similarity search, composition, and downstream conditioning through simple vector-space operations. Finally, we demonstrate that SLS generalizes across diverse tasks reducing their FLOPs by over 150 times compared to token-based approaches, and establishing a general-purpose latent foundation for vector graphics research.
CommentsAccepted at The 19th European Conference on Computer Vision -- ECCV 2026