发表机构
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ); School of Integrated Circuits (School of Information Science and Electronic Engineering), Shanghai Jiao Tong University; Zhiyuan College, Shanghai Jiao Tong University(广东省人工智能与数字经济实验室(深圳); 上海交通大学集成电路学院(信息科学与工程学院); 上海交通大学致远学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对文本到SVG生成评估的缺陷,提出视觉基础的SVGEval基准,发现多模态模型在几何布局判断上的差距,训练出可解释的SVG质量评分器,为SVG生成评估改进提供支撑。
AI 中文摘要
多模态大模型越来越多地被用于生成可缩放矢量图形(SVG),但可靠的评估方法仍有待深入探索。现有评估协议通常以代码为中心,或在渲染SVG后借用光栅图像指标,无法反映人类感知,且忽略了SVG特有的属性,如几何特性和空间构图。我们提出SVGEval,这是一个以视觉为基础的多模态基准,用于与人类对齐的SVG质量评估。SVGEval明确结合视觉渲染结果,以评估模型能否判断渲染结果,而非仅检查SVG代码,还提供了经多轮人工标注并由专家优化的高质量注释。对代表性多模态模型的系统评估显示存在明显差距:模型在语义对齐和美学方面表现相对较好,但在与几何和布局相关的判断上表现不佳。基于SVGEval,我们训练了一个可解释的SVG质量评分器,可输出多维度分数及文本形式的理由。消融实验表明,明确的视觉基础和推理监督至关重要,尤其对空间和几何评估而言。SVGEval为多模态模型时代SVG生成的评估与改进提供了可靠的测试平台和实用的评分器。
英文摘要
Multimodal large models are increasingly used to generate scalable vector graphics (SVG), but reliable evaluation remains underexplored. Existing protocols are often code-centric or borrow raster-image metrics after rendering SVGs, which fail to reflect human perception and overlook SVG-specific qualities such as geometry and spatial composition. We introduce SVGEval, a vision-grounded multimodal benchmark for human-aligned SVG quality assessment. SVGEval explicitly incorporates visual renderings to evaluate whether models can judge the rendered outcome rather than only inspect SVG code, and provides high-quality annotations obtained via multi-round human labeling with expert refinement. Systematic evaluations across representative multimodal models reveal a clear gap: models perform relatively well on semantic alignment and aesthetics, yet struggle on geometry- and layout-related judgments. Building on SVGEval, we train an explainable SVG quality scorer that outputs multi-aspect scores with textual rationales. Ablations show that explicit visual grounding and reasoning supervision are crucial, especially for spatial and geometric assessment. SVGEval offers a reliable testbed and practical scorer for evaluating and improving SVG generation in the era of multimodal models.
CommentsAccepted by ECCV 2026