发表机构
Cardiff University; The Cyprus Institute; Carleton University; University of Sheffield(卡迪夫大学; 塞浦路斯研究所; 卡尔顿大学; 谢菲尔德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对风格迁移缺乏可靠评估标准的问题,本文提出ASTRA框架,包含基准数据ASTRA-Data和学习评估器ASTRA-Score,通过两阶段成对比较获取人类偏好,实现与人类排名高度相关的自动评估。
AI 中文摘要
风格迁移缺乏可靠的评估标准:真实标签本质上是不明确的,现有的自动评估指标往往无法反映人类偏好。本文提出了ASTRA(风格迁移算法评估),一种用于自动评估风格迁移算法的方法;它包含两个组成部分,ASTRA-Data和ASTRA-Score。ASTRA-Data由一个包含内容参考和风格参考的基准图像集、在该基准集上生成的风格迁移结果集合,以及通过两阶段成对比较协议捕获人类判断的用户研究数据组成。从这些标注中,我们推导出基于排名的内容保留、风格保真度和整体偏好的真实标签。基于ASTRA-Data,我们构建了ASTRA-Score,一个学习型评估器,它从内容-风格-风格化图像三元组中预测与偏好对齐的分数,从而能够对应用于基准集的新模型进行自动且可扩展的评估。实验结果表明,与先前的指标相比,ASTRA-Score与人类排名的相关性显著更高。总体而言,ASTRA建立了一种用于风格迁移方法标准化评估的稳健机制。
英文摘要
Style transfer lacks a reliable evaluation standard: ground truth is inherently ill-defined, and existing automatic metrics often fail to reflect human preference. This paper introduces ASTRA (Assessment of Style TRansfer Algorithms), an approach for automatic evaluation of style transfer algorithms; it contains two components, ASTRA-Data and ASTRA-Score. ASTRA-Data consists of a benchmark image set of content and style references, a collection of style transfer results generated on the benchmark set, and user study data capturing human judgements through a two-stage pairwise comparison protocol. From these annotations, we derive ranking-based ground truth for content preservation, style fidelity, and overall preference. Based on ASTRA-Data, we construct ASTRA-Score, a learnt evaluator that predicts preference-aligned scores from content-style-stylization image triplets, enabling automatic and scalable evaluation of new models applied to the benchmark set. Experimental results demonstrate that ASTRA-Score achieves substantially higher correlation with human rankings compared to prior metrics. Overall, ASTRA establishes a robust mechanism for standardised evaluation of style transfer methods.