SketchFlow:基于CLIP潜在空间中GMM先验流的零样本矢量草图生成
SketchFlow: Zero-Shot Vector Sketch Generation via GMM Prior Flow in CLIP Latent Space
查看机构详情
- Guangdong Provincial Key Laboratory of Visual Media and Multidimensional Intelligence(广东省视觉媒体与多维智能重点实验室)
- Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
SketchFlow是基于CLIP潜在空间的零样本矢量草图生成框架,通过GMM先验流匹配与混合扩散解码器,实现了优于基线的视觉质量与零样本泛化能力。
中文摘要 AI 辅助
矢量草图是人类进行抽象表达最简洁、直观的媒介之一。然而,由于细粒度、高质量的文本-草图配对数据极度匮乏,生成具有人类手绘风格的高质量矢量笔触仍是一个未解决的挑战。现有的文本条件生成方法通常依赖不稳定、耗时的优化,或难以以零样本方式泛化到未见类别。为解决这些局限,我们提出SketchFlow,一个基于最优传输(OT)理论和流匹配的新型生成框架。通过利用预训练的CLIP模型绕过繁琐的图像级文本标注,我们将跨模态对齐直接在CLIP潜在空间中表述为连续映射问题。为弥合离散文本概念与连续草图特征之间不可避免的模态差距,我们首先向离散类别嵌入注入噪声,以构建连续高斯混合模型(GMM)先验。随后,我们使用最优传输条件流匹配(OT-CFM)模型学习确定性向量场,将该连续GMM先验映射到目标草图特征分布。最后,我们设计了融合1D U-Net和Transformer架构的混合扩散解码器,以将这些特征解码为快速且高保真的笔触轨迹。大量实验表明,SketchFlow在视觉质量和对自然人类手绘风格的贴合度上显著优于现有基线。此外,我们的保几何框架对QuickDraw训练词汇之外的提示(包括未见概念标签和语义修饰词)展现出良好的局部零样本合成能力,同时支持不同概念间平滑、连续的语义插值。源代码可在:this https URL获取。
英文摘要
Vector sketches remain one of the most concise and immediate mediums for abstract human expression. However, generating high-quality vector strokes that exhibit human-like drawing styles remains an open challenge due to the severe scarcity of fine-grained, high-quality text-to-sketch paired data. Existing text-conditioned generation methods often rely on unstable, time-consuming optimization or struggle to generalize to unseen categories in a zero-shot manner. To address these limitations, we present SketchFlow, a novel generative framework rooted in Optimal Transport (OT) theory and flow matching. By leveraging pre-trained CLIP models to bypass labor-intensive image-level text annotations, we formulate cross-modal alignment as a continuous mapping problem directly within the CLIP latent space. To bridge the inevitable modality gap between discrete text concepts and continuous sketch features, we first inject noise into discrete category embeddings to construct a continuous Gaussian Mixture Model (GMM) prior. We then utilize an Optimal Transport Conditional Flow Matching (OT-CFM) model to learn a deterministic vector field mapping from this continuous GMM prior to the target sketch feature distribution. Finally, a Hybrid Diffusion Decoder, fusing 1D U-Net and Transformer architectures, is designed to decode these features into fast and high-fidelity stroke trajectories. Extensive experiments demonstrate that SketchFlow substantially outperforms existing baselines in visual quality and adherence to natural human drawing styles. Furthermore, our geometry-preserving framework demonstrates promising local zero-shot synthesis for prompts beyond the QuickDraw training vocabulary, including unseen concept labels and semantic modifiers, while enabling smooth, continuous semantic interpolation between distinct concepts. Source code is available at: https://github.com/doudin404/SketchFlow.