风格还是签名?冻结视觉嵌入中风格分类的艺术家不相交评估
Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings
浏览论文内容
中文总结 AI 辅助
该研究提出艺术家不相交评估协议,发现CLIP等模型的冻结图像嵌入风格分类准确率在该协议下下降,超现实主义降幅最大,验证了冻结嵌入中风格理解需此类评估。
中文摘要 AI 辅助
CLIP等模型生成的冻结图像嵌入被越来越多地用于按艺术史风格对绘画进行分类,且报告的准确率很高。我们询问这种准确率是反映了对风格的理解,还是对个别艺术家的识别。标准评估采用随机划分,使得同一艺术家的作品同时出现在训练集和测试集中,因此分类器可以通过识别画家而非艺术流派取得成功。我们在艺术家不相交协议下重新评估风格分类,每次留出一位艺术家的所有作品,确保任何作品都不会使用其所属画家的其他作品进行分类。在包含四个20世纪流派的320幅绘画的平衡数据集上,5-近邻(5-NN)风格准确率在该协议下从0.87降至0.77,且下降幅度极不均衡:印象派和立体派几乎没有变化,而超现实主义下降了20个百分点。该模式在包括仅视觉自监督模型在内的四个图像编码器中均成立,表明该效果源于视觉结构而非语言。当编码器捕获到真实的共享形式时,个别艺术家几乎无法被识别,但风格却保持稳健;而超现实主义则呈现相反情况。我们认为,艺术家不相交评估对于测量冻结嵌入中的风格理解是必要的。
英文摘要
Frozen image embeddings from models such as CLIP are increasingly used to classify paintings by art-historical style, with high reported accuracy. We ask whether this accuracy reflects an understanding of style or the recognition of individual artists. Standard evaluation uses random splits in which works by the same artist appear on both sides, so a classifier can succeed by recognising the painter rather than the movement. We re-evaluate style classification under an artist-disjoint protocol, holding out every artist in turn so that no work is ever classified using other works by its own painter. On a balanced dataset of 320 paintings across four twentieth-century movements, 5-NN style accuracy falls from 0.87 to 0.77 under this protocol, and the drop is sharply uneven. Impressionism and Cubism barely move, while Surrealism falls twenty points. The pattern holds across four image encoders, including a vision-only self-supervised model, which places the effect in visual structure rather than language. Where an encoder captures genuine shared form, individual artists are barely recognisable yet style is robust, while Surrealism shows the opposite. We argue that artist-disjoint evaluation is necessary to measure stylistic understanding in frozen embeddings.
发表机构
- University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。