arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于零样本SAM2分割与视觉Transformer的破损泥板埃兰楔形文字符号识别

Zero-Shot SAM2 Segmentation and Vision Transformer-Based Recognition of Elamite Cuneiform Symbols from Degraded Tablet Images

Utsav Poudel, Rasik Bhattarai, Siddhartha Pathak, Raghavendra Ramacharna, Gaurav Jaswal

arXiv 2608.18544首次发表:更新:

发表机构

Vellore Institute of Technology; Indian Institute of Technology; WestCliff University; York St John University; Norwegian University of Science and Technology(韦洛尔理工学院; 印度理工学院; 韦斯特克利夫大学; 约克圣约翰大学; 挪威科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对破损泥板埃兰楔形文字识别的信号退化与类别不平衡问题,提出EpigraphNet分割引导Transformer流水线,在基准测试中准确率优于多种主流模型,实现高效均衡的符号识别。

AI 中文摘要

古代楔形文字的自动识别面临复杂的信号退化问题:黏土泥板的三维浮雕结构会产生空间变化的光照和阴影,表面侵蚀引入的结构化噪声与真实符号印记重叠,且141个符号类别间存在严重的类别不平衡,削弱了分类器的可靠性。本文提出EpigraphNet,一种分割引导的Transformer流水线,在波斯波利斯 fortification档案上进行评估。从1239张标注泥板图像出发,采用亮度自适应形态学预处理和零样本SAM2-Large分割生成干净的二值符号掩码,再由经过逆频率类别权重微调的Vision Transformer(ViT-B/16)进行分类。EpigraphNet在132类基准测试中达到86.41%的top-1准确率,比最强的CNN基线(ResNet-101,69.20%)提升17.21个百分点,在相同条件下比四种现代骨干网络(DeiT-B/16、Swin-B、ConvNeXt-B、EfficientNet-B4)提升5.31%-12.91%。该流水线在NVIDIA A100 GPU上每个符号运行约18毫秒,符号频率与每类性能之间较低的斯皮尔曼相关性表明,频繁和稀有类别的识别更均衡。实现代码可在指定网址获取。

英文摘要

Automated recognition of ancient cuneiform script poses a compound signal-degradation problem: the three-dimensional relief of clay tablets creates spatially varying illumination and cast shadows, surface erosion introduces structured noise that overlaps with genuine sign impressions, and severe class imbalance across 141 sign categories undermines classifier reliability. We introduce EpigraphNet, a segmentation-guided transformer pipeline evaluated on the Persepolis Fortification Archive. From 1,239 annotated tablet images, brightness-adaptive morphological preprocessing and zero-shot SAM2-Large segmentation generate clean binary symbol masks, which a fine-tuned Vision Transformer (ViT-B/16) with inverse-frequency class weighting then classifies. EpigraphNet reaches 86.41% top-1 accuracy on a 132-class benchmark, a 17.21 percentage-point gain over the strongest CNN baseline (ResNet-101, 69.20%) and 5.31-12.91% over four modern backbones (DeiT-B/16, Swin-B, ConvNeXt-B, EfficientNet-B4) under identical conditions. The full pipeline runs at approximately 18 ms per sign on an NVIDIA A100 GPU. A lower Spearman correlation between sign frequency and per-class performance indicates more balanced recognition across frequent and rare classes. Implementation is available at: github.com/r11up/sam-guided-vit

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑