发表机构
SB Intuitions; University of Tokyo(SB Intuitions; 东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对细粒度图像编辑缺乏可验证评估的问题,提出VeriEdit-Bench基准,基于SVG等结构化资产提供确定性四轴评分,揭示模型能力与失败模式。
AI 中文摘要
细粒度图像编辑需要的不仅仅是产生视觉上合理的结果:编辑器必须精确地执行所请求的属性更改,同时保持其他所有内容不变。然而,现有的基准在真实性和可验证性之间留下了一个关键缺口:基于真实图像的基准通常依赖人类或视觉-语言模型的判断,而确定性评估主要集中于合成形状画布,面向应用的扩展主要限于图表。这使得难以精确确定所请求的编辑被执行了多少、意外更改发生在何处,以及模型之间的微小差异是反映了真实的编辑能力还是评估者的不确定性。为弥合这一缺口,我们提出了VeriEdit-Bench,一个针对现实结构化资产的细粒度、指令忠实图像编辑基准,具有确定性的四轴评估。其1,740个案例来自153个可缩放矢量图形(SVG)图形、图表、Web界面和演示幻灯片的源代码。受控的源代码编辑保留了原始视觉上下文,同时产生了精确的目标图像、像素级编辑掩码和明确的编辑规范,使得沿四个轴进行可复现的评分成为可能:编辑保真度、保留度、定位和幅度。评估了十一个编辑器,我们发现即使是最强的模型也远未达到满分;对于相同的重新着色操作,排名在图表和SVG图形之间发生反转;具有相似像素精度特征的输出在定位和更改幅度上仍可能存在显著差异。这种分解产生了分级的、可验证的反馈,并揭示了整体分数或依赖评估者的判断可能掩盖的模型特定能力和失败模式。
英文摘要
Fine-grained image editing requires more than producing a visually plausible result: an editor must execute the requested attribute change precisely while leaving everything else intact. However, existing benchmarks leave a critical gap between realism and verifiability: benchmarks built on realistic images typically rely on human or vision--language model judgments, while deterministic evaluation has largely focused on synthetic shape canvases, with application-oriented extensions primarily limited to charts. This makes it difficult to determine precisely how much of a requested edit was executed, where unintended changes occurred, and whether small differences between models reflect genuine editing capability or evaluator uncertainty. To bridge this gap, we present VeriEdit-Bench, a benchmark for fine-grained, instruction-faithful image editing across realistic structured assets with deterministic, four-axis evaluation. Its 1,740 cases are compiled from the source code of 153 Scalable Vector Graphics (SVG) graphics, charts, web interfaces, and presentation slides. Controlled source-code edits preserve the original visual context while yielding exact target images, pixel-level edit masks, and explicit edit specifications, enabling reproducible scoring along four axes: edit fidelity, preservation, localization, and magnitude. Evaluating eleven editors, we find that even the strongest model remains far from full credit; rankings for the same recoloring operation reverse between charts and SVG graphics; and outputs with similar pixel-accuracy profiles can still differ substantially in localization and change magnitude. This decomposition yields graded, verifiable feedback and exposes model-specific capability and failure profiles that holistic scores or evaluator-dependent judgments may obscure.
CommentsImage editing benchmark