arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12088cs.CV

RA-ClipScore:让生成模型评估更具可解释性

RA-ClipScore: Making Generative Model Evaluation More Interpretable

发表机构皇家理工学院 · 瑞典国家图书馆 · 美国艺电公司
另 1 家 · 查看机构详情
  • KTH Royal Institute of Technology(皇家理工学院)
  • National Library of Sweden(瑞典国家图书馆)
  • Electronic Arts (EA)(美国艺电公司)
  • SEED

机构由 AI 辅助整理,请以论文原文为准。

Yifan Lu, Taras Kucherenko, Hedvig Kjellström, Judith Bütepage

首次发表
浏览论文内容

中文总结 AI 辅助

RA-CLIPScore是一种新型生成模型评估指标,通过双提示与局部补丁标记实现属性级和空间分布对齐评估,比现有方法更鲁棒可解释,能揭示生成模型的空间偏差,更贴合人类视觉多样性感知。

中文摘要 AI 辅助

生成模型能够产生与真实数据几乎难以区分的图像,但严格且可解释的评估仍然具有挑战性。传统指标如FID仅提供标量分数,诊断性见解有限;广泛采用的基于CLIP的指标能够实现超越简单训练类标签的语义评估,但继承了CLIP训练范式的限制,阻碍了属性级分析。我们提出RA-CLIPScore,一种新的指标,可缓解这些问题并将基于CLIP的评估扩展到空间分布对齐,衡量生成对象是否符合训练数据中的位置先验。RA-CLIPScore引入双提示以解耦竞争属性,并利用局部补丁标记捕捉细粒度区域语义。我们评估图像生成模型匹配训练数据的属性和空间分布的能力,大量实验表明,RA-CLIPScore比现有方法提供更鲁棒、可解释的评估,尤其在分布不对齐或部分不相关文本属性的情况下;我们进一步证明它如何揭示生成模型中的空间偏差,用户评估证实,基于RA-CLIPScore的区域单属性分歧与人类对视觉多样性的感知比现有语义指标更一致。

英文摘要

Generative models can produce images nearly indistinguishable from real data, yet rigorous and interpretable evaluation remains challenging. Conventional metrics such as FID provide only scalar scores with limited diagnostic insight. Widely adopted CLIP-based metrics enable semantic evaluation beyond simple training class labels, but inherit limitations from CLIP's training paradigm that restrict attribute-wise analysis. We propose RA-CLIPScore, a novel metric that mitigates these issues and extends CLIP-based evaluation to spatial distribution alignment, measuring whether generated objects adhere to the positional priors found in the training data. RA-CLIPScore introduces dual prompts to decouple competing attributes and leverages local patch tokens to capture fine-grained regional semantics. We evaluate image generative models on their ability to match both attribute and spatial distributions of the training data. Extensive experiments show that RA-CLIPScore provides more robust and interpretable evaluations than prior methods, particularly under distribution misalignment or partially irrelevant textual attributes. We further demonstrate how it reveals spatial biases in generative models. User evaluations confirm that Regional Single Attribute Divergence based on our RA-CLIPScore aligns more closely with human perception of visual diversity than existing semantic metrics.

↑