arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03806cs.AIcs.CV

SVG-Score:面向文本生成可缩放矢量图形(SVG)的人类对齐评估

SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation

Marco Cipriano, Leonardo Zini, Alexandra Schild, Valentin Teutschbein, Afsana Mimi, Marcella Cornia, Lorenzo Baraldi, Gerard de Melo

首次发表
浏览论文内容

中文总结 AI 辅助

针对文本生成SVG缺乏专用评估指标的问题,提出人类对齐的SVG-Scope框架,开发两类互补评估器并对主流SVG生成器开展基准测试。

中文摘要 AI 辅助

随着生成模型在表达能力和可控性上的提升,可缩放矢量图形(SVG)生成正受到越来越多的关注。然而,该领域的进展因缺乏专用评估协议而受阻:当前实践依赖为自然图像设计的指标,最具代表性的是CLIPScore,该指标从未在矢量图形上进行训练,且仅与人类判断部分对齐。我们提出了\textbf{\textsc{ours}},一个面向文本生成SVG的人类对齐评估框架。通过受控的描述和图像扰动,我们首先证明基于CLIP的指标几乎不会对SVG生成器实际产生的错误(如颜色错误、数量错误、空间关系错误)做出反应,而现成的视觉语言模型(VLM)评估器虽更敏感,但对不同错误类型和SVG风格的反应不一致。随后,我们引入了一个用于\textit{语义对齐}的人工标注数据集,用于衡量生成的SVG对其描述的忠实度。基于该数据集,我们开发了两个互补的评估器:适配矢量图形并对齐人类偏好的CLIP评分器,用于快速大规模评估;以及通过监督微调与奖励型强化学习训练的VLM评估器,用于更具表达性和可解释性的评估。利用这两个评估器,我们在独立的描述集上对主要的开源、商业及基于优化的SVG生成器进行了基准测试。

英文摘要

Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in expressiveness and controllability. Progress, however, is held back by the lack of domain-specific evaluation protocols: current practice relies on metrics designed for natural images, most notably CLIPScore, which was never trained on vector graphics and aligns only partially with human judgment. We introduce \textbf{\ours}, a human-aligned evaluation framework for text-to-SVG generation. Through controlled caption and image perturbations, we first show that CLIP-based scores barely react to the errors SVG generators actually make, such as wrong colors, counts, and spatial relations, and that off-the-shelf Vision-Language Model (VLM) judges, while more sensitive, respond unevenly across error types and SVG styles. We then introduce a human-annotated dataset for \textit{Semantic Alignment}, measuring how faithfully a generated SVG reflects its caption. Building on it, we develop two complementary evaluators: CLIP scorers adapted to vector graphics and then aligned to human preferences, for fast large-scale evaluation, and a VLM judge trained with supervised fine-tuning and reward-shaped reinforcement learning, for more expressive and interpretable assessment. Using both, we benchmark major open-source, commercial, and optimization-based SVG generators on an independent caption set.

发表机构

  • Hasso-Plattner Institute(哈索·普拉特纳研究所)
  • University of Modena and Reggio Emilia(摩德纳-雷焦艾米利亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑