SARCLIP:用于17世纪西班牙美洲公证记录的可扩展CLIP检索系统
SARCLIP: A Scalable CLIP-Based Retrieval System for Seventeenth-Century Spanish American Notary Records
浏览论文内容
中文总结 AI 辅助
该研究针对历史手稿档案检索难题,构建了基于CLIP的SARCLIP系统,结合FAISS索引、伪相关反馈等人在回路功能,可在大规模17世纪西班牙美洲公证记录语料库上实现完整交互式检索。
中文摘要 AI 辅助
历史手稿档案因笔迹不一致、拼写古旧且缺乏大规模可靠转录,难以进行标准文本检索。我们提出SARCLIP(西班牙美洲公证记录与CLIP结合系统),这是为阿根廷国家档案馆17世纪西班牙美洲公证记录构建的已部署检索系统,该语料库包含超过1360万个词-图像补丁,涵盖100多卷缩微胶卷(“rollos”)。SARCLIP基于CLIP ViT-B/16模型构建,该模型经古文字学专家标注数据进行对比微调,且在三方面扩展了现有工作:(1)通过FAISS索引将近似最近邻检索扩展至近完整语料库;(2)通过伪相关反馈(Rocchio)优化top-k结果;(3)通过可视化文档浏览、基于画布的补丁标注及定期模型再训练构建人在回路的循环。与仅在5卷小型子集上评估检索效果的初始研究原型不同,SARCLIP被证实是可在近完整语料库上运行的完整交互式工具,演示期间参会者可现场体验完整的搜索、浏览、标注及再训练工作流程。
英文摘要
Historical manuscript archives resist standard text search due to inconsistent handwriting, archaic orthography, and the absence of reliable transcriptions at scale. We present SARCLIP (Spanish American Notary Records Meets CLIP), a deployed retrieval system for the National Archives of Argentina's seventeenth-century Spanish American notary records, a corpus of more than 13.6 million word-image patches spanning over 100 microfilm rolls ("rollos"). SARCLIP is built on a CLIP ViT-B/16 model contrastively fine-tuned on paleography-expert-annotated data, and extends prior work by (1) scaling approximate nearest-neighbor retrieval to the near-complete corpus via a FAISS index, (2) refining top-k results through pseudo-relevance feedback (Rocchio), and (3) closing a human-in-the-loop cycle through visual document browsing, canvas-based patch annotation, and periodic model retraining. Unlike the system's initial research prototype, which evaluated retrieval on a small five-rollo subset, SARCLIP is demonstrated as a complete, interactive tool operating over the near-complete corpus. Attendees experience the full search, browse, annotate, and retrain workflow live during this demonstration.