历史手稿插图视觉分类数据集与模型评估
A Visual Classification Dataset and Model Evaluation for Historical Manuscript Illustrations
- Ben-Gurion University of the Negev(内盖夫本-古里安大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究构建了含15,000幅历史手稿插图的22类数据集,评估多种视觉模型,发现微调分类器最优,ConvNeXt达88.9%准确率,并指出通用模型局限及分类法影响。
AI中文摘要:
历史手稿插图保存了丰富的过去文化视觉证据。它们描绘了人物、动物、植物、图表、音乐记谱和装饰形式。尽管大型数字化项目已使许多手稿在线可获取,但该材料本身仍难以大规模探索。提取系统可以在手稿页面上找到插图,但若没有有意义的类别,大型收藏仍难以搜索和探索。我们通过引入一个包含来自数百年前手稿的15,000幅插图、涵盖22个类别的人工标注数据集,并在此任务上评估现代视觉模型用于图像分类,来解决这一空白。由于风格多样性、退化以及语义模糊性,该问题具有挑战性,许多图像符合多个类别。我们比较了微调的CNN和基于Transformer的分类器、零样本CLIP、基于嵌入的分类器以及直接视觉-语言模型。结果表明,微调图像分类器总体表现最佳,其中ConvNeXt达到88.9%的准确率和81.3%的宏F1分数。使用CLIP嵌入结合XGBoost提供了一个强有力的替代方案。相比之下,零样本CLIP和直接视觉-语言分类表现明显较差,凸显了通用模型在该领域的局限性。除整体性能外,分析还揭示了哪些类别在视觉上可分离,以及哪些错误反映了真实的语义重叠,表明一些局限性源于分类法本身。
英文摘要:
Historical manuscript illustrations preserve rich visual evidence of past cultures. They depict people, animals, plants, diagrams, music notations, and decorative forms. Although large digitization projects have made many manuscripts available online, the material itself remains difficult to explore at scale. Extraction systems can find illustrations on manuscript pages, but without meaningful categories, large collections remain hard to search and explore. We address this gap by introducing a manually labeled dataset of 15,000 illustrations from manuscripts dating back hundreds of years across 22 categories, and evaluating modern vision models for image classification on this task. The problem is challenging due to stylistic diversity, degradation, and semantic ambiguity, with many images that fit more than one category. We compare fine-tuned CNN and Transformer-based classifiers, zero-shot CLIP, embedding-based classifiers, and direct vision-language models. Results show that fine-tuned image classifiers perform best overall, with ConvNeXt reaching 88.9% accuracy and 81.3% macro-F1. Using CLIP embeddings with XGBoost provides a strong alternative. In contrast, zero-shot CLIP and direct vision-language classification perform substantially worse, highlighting the limits of general-purpose models in this domain. Beyond overall performance, the analysis reveals which categories are visually separable and where errors reflect genuine semantic overlap, suggesting that some limitations arise from the taxonomy itself.