发表机构
Computer Vision Center (CVC); Universitat Autònoma de Barcelona; Stockholm University(计算机视觉中心; 巴塞罗那自治大学; 斯德哥尔摩大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出三阶段无监督域适应流水线,结合SimCLR+DANN编码器和嵌入空间风格适应,用于历史加密手稿符号识别,在七份手稿集上大幅超越CLIP和DINOv2等基线。
AI 中文摘要
历史加密手稿的解密在数字人文学科中构成了一项根本性挑战:在任何转录开始之前,必须首先识别并描述底层密码字母表的符号清单。我们通过符号识别来解决这一挑战:给定一个以一组渲染字体字形形式指定的候选字母表,任务是确定其字符是否以及出现在何处,出现在未见过的手写文档中,且无需目标脚本的任何标注示例。主要困难在于干净的数字渲染字体查询与退化手写手稿符号之间的域差距。我们提出了一种无需手动标注即可弥合这一差距的三阶段流水线,将联合SimCLR+DANN编码器用于域不变字形表示,并在检索时应用嵌入空间风格适应机制,无需重新训练。在来自七个加密手稿收藏的十四页上的实验表明,我们的方法大幅优于零样本基础模型(包括CLIP和DINOv2),在P@1上比CLIP ViT-L/14高出+0.194,并比任务特定训练基线高出+0.138 P@1。我们进一步证明,在完全无监督设置下计算的Raw-Cover指标提供了一种有意义的脚本族指纹,可识别未知文档的底层字母表。这一能力对古文字学家、历史学家以及其他研究未解密手稿的研究人员具有直接的实际相关性。
英文摘要
The decipherment of historical encrypted manuscripts poses a fundamental challenge in Digital Humanities: before any transcription can begin, the symbol inventory of the underlying cipher alphabet must first be identified and characterized. We address this challenge through symbol spotting: given a candidate alphabet specified as a set of rendered font glyphs, the task is to determine whether and where its characters appear in an unseen handwritten document, without any labeled examples from the target script. The main difficulty lies in the domain gap between clean, digitally rendered font queries and degraded handwritten manuscript symbols. We propose a three-stage pipeline that bridges this gap without manual annotation, combining a joint SimCLR+DANN encoder for domain-invariant glyph representations with an embedding-space style-adaptation mechanism applied at retrieval time, requiring no re-training. Experiments on fourteen pages from seven encrypted manuscript collections show that our method outperforms zero-shot foundation models, including CLIP and DINOv2, by a large margin ($+0.194$ P@1 over CLIP ViT-L/14), and surpasses task-specific trained baselines by $+0.138$ P@1. We further demonstrate that the Raw-Cover metric, computed in a fully unsupervised setting, provides a meaningful script-family fingerprint that identifies the underlying alphabet of an unknown document. This capability is of direct practical relevance to palaeographers, historians, and other researchers working with undeciphered manuscripts.