发表机构
Adam Mickiewicz University; ArtiCollect; KAUST; Kiel University(亚当·密茨凯维奇大学; 艺术收藏机构; 沙特阿卜杜拉国王科技大学; 基尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对实例级艺术品识别难题,提出SynGallery合成图库数据集。从真实绘画目录图像出发,经3D场景渲染生成数据集。实验表明其合成视图训练信号强,能提升识别率,增益主要源于几何视角变化。
AI 中文摘要
实例级艺术品识别需要将手持游客照片与大型博物馆收藏中的特定作品进行匹配。这具有挑战性,因为绘画数据集通常提供干净的目录图像用于训练,而测试查询是在倾斜视角、画廊灯光、反射、画框和其他场景级变化下拍摄的。我们提出了SynGallery,一个用于艺术品检索的合成图库数据集,无需收集额外的真实照片即可解决这一差距。从真实绘画的目录图像开始,我们将每件艺术品放置在程序生成的3D画廊场景中,并在各种几何和外观条件下从多个视角进行渲染,同时保留原始作品的准确身份。生成的数据集包含来自大都会基准的4898幅绘画的24490个渲染视图。我们表明,这些合成视图提供了比相应工作室照片更强的训练信号。在相同数量的训练数据点下,仅在SynGallery上训练可将艺术绘画识别率从67.18提高到73.47 GAP$^-$。当添加到完整的大都会训练集中时,SynGallery将已发布的基准协议从35.97提高到38.48 GAP。消融实验表明,增益主要来自几何视角变化而非摄影逼真度:模糊、传感器噪声和图像压缩会持续降低性能。
英文摘要
Instance-level artwork recognition requires matching a handheld visitor photograph to a specific work in a large museum collection. This is challenging because painting datasets typically provide clean catalog images for training, while test queries are captured under oblique viewpoints, gallery lighting, reflections, frames, and other scene-level variations. We present SynGallery, a synthetic gallery dataset for artwork retrieval that addresses this gap without collecting additional real photographs. Starting from catalog images of real paintings, we place each artwork into a procedurally generated 3D gallery scene and render it from multiple viewpoints under varied geometric and appearance conditions, while preserving the exact identity of the original work. The resulting dataset contains 24,490 rendered views of 4,898 paintings from the Met benchmark. We show that these synthetic views provide a stronger training signal than the corresponding studio photographs. At the same number of training data points, training only on SynGallery improves art painting recognition from 67.18 to 73.47 GAP$^-$. When added to the full Met training set, SynGallery improves the published benchmark protocol from 35.97 to 38.48 GAP. Ablation experiments show that the gain comes from scene-level view variation rather than photographic realism: reducing the five rendered viewpoints to a single frontal view removes most of the improvement, while simulating capture artifacts such as blur, sensor noise, and image compression consistently reduces performance.