向量基准测试:模型能否精准编辑 SVG 代码?
Vector-Bench: Can Models Surgically Edit SVG Code?
查看机构详情
- Theta Labs(西塔实验室)
- Warping(沃平)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究向量编辑中模型能否精准编辑 SVG 代码的问题,引入 Vector-Bench 基准测试,含 40 个 SVG 修复任务,定义规范奖励,评估 34 个模型端点,发现最强端点完整规范成功率仅 15.0%,揭示修复进度与精准编辑差异大并公开相关数据。
中文摘要 AI 辅助
基于指令的向量编辑需要两项能力:进行所需更改并保持其他部分不变。当仅将输出判断为光栅图像时,第二项能力容易被忽略。我们引入了 Vector-Bench,这是一个包含 40 个 SVG 修复任务的紧凑且具有挑战性的基准测试。每个任务将损坏的 SVG 程序与作者编写的视觉指令、隐藏的目标程序、平均 5.05 个带注释的修复以及平均 60.55 个受保护对象配对。指令描述可见缺陷而不暴露元素标识符、坐标、颜色代码或路径数据。我们定义了一个确定性二元规范奖励:所需修复使用属性感知感知容差,而未请求的渲染或应用相关结构必须在语义上保持不变,并且结果必须是有效的 SVG。保留规范目标相等性和更严格的源保真度作为诊断。有效性门控修复进度、近乎完整的层级和有效输出意外更改率(UCR)解释部分结果。我们在 1360 个请求上评估了 34 个模型端点(25 个列为开放权重、5 个低成本控件和 4 个前沿封闭端点)。尽管平均修复进度为 43.7%,但最强的端点仅达到 15.0%的完整规范成功率,这表明明显的修复进度和规范忠实编辑仍然有很大差异。所有提示、输出、评分代码、成本和每个任务的报告都已发布。
英文摘要
Instruction-based vector editing requires two capabilities: making a requested change and leaving everything else alone. The second is easy to miss when an output is judged only as a raster image. We introduce Vector-Bench, a compact, difficult benchmark of 40 SVG repair tasks. Each task pairs a corrupted SVG program with an author-written visual instruction, a hidden target program, 5.05 annotated repairs on average, and an average of 60.55 protected objects. Instructions describe visible defects without exposing element identifiers, coordinates, color codes, or path data. We define a deterministic binary specification reward: requested repairs use attribute-aware perceptual tolerances, while unrequested rendering- or application-relevant structure must remain semantically unchanged and the result must be a valid SVG. Canonical target equality and stricter source fidelity are retained as diagnostics. Validity-gated repair progress, a near-complete tier, and valid-output Unintended Change Rate (UCR) explain partial outcomes. We evaluate 34 model endpoints (25 listed as open-weight, 5 inexpensive controls, and 4 frontier closed endpoints) over 1360 requests. The strongest endpoint reaches only 15.0% full specification success, despite 43.7% mean repair progress, showing that apparent repair progress and specification-faithful editing remain substantially different. All prompts, outputs, scoring code, costs, and per-task reports are released.