发表机构
University of Graz; University of Basel; Università degli Studi di Napoli Federico II(格拉茨大学; 巴塞尔大学; 那不勒斯费德里科二世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究中世纪文本字符集切换问题,用一对一及带状循环神经网络训练字符映射,可恢复字符错误率,还能扩展缩写,引入字母词形还原启发式方法并展示Python库来执行相关方法。
AI 中文摘要
中世纪文档抄写员的做法差异很大,再加上异构数字化政策,导致语料库中的字符集必须被视为可变的。本文旨在解决灵活切换字符集的问题。我们专注于一对一的字符映射,并训练字符级的一对一循环神经网络进行自我监督以撤销这些映射,即使只有20行文本也能恢复一半的字符错误率。我们分析了这些一对一网络在光学字符识别后校正中的应用,发现它们在完全忽略插入和删除的情况下仍有显著改进。然后,我们使用从平行语料库编译的具有字符级对齐真值的完全相同的网络,以一种我们称为带状循环神经网络的训练和推理模式成功扩展了中世纪宪章转录中的缩写。最后,我们引入了一种精心设计的启发式方法,该方法采用任意两个字符集的字符,并定义了一个度量来封装我们认为的字符语义相似性。我们将这种映射的构建称为字母词形还原,并展示了一个丰富的Python库,该库能高效地执行所有提出的方法。
英文摘要
Medieval document transcribers have very different practices; on top of that, heterogeneous digitization policies have resulted in corpora where the character-set must be viewed as fluid. In this paper we address the problem of changing between character-sets in a flexible manner. We focus on one-to-one character mappings and train characterlevel one-to-one RNNs to undo them with self-supervision; recovering half the CER even with 20 text lines. We analyse the use of these one-to-one networks for HTR post-correction and we see that they obtain significant improvements while totally ignoring ins-dels. We then use the exact same networks with character-level alignment groundtruth compiled from parallel corpora in a training and inference mode we call Banded RNNs. We use such networks to successfully expand abbreviations in medieval charter transcriptions. Finally we introduce an elaborate heuristic which takes the characters of two arbitrary character-sets and defines a metric encapsulating what we consider to be semantic similarity of characters. We call the construction of such mappings letter lemmatization and present a rich Python library that efficiently performs all presented methods.
CommentsAccepted for publication (after peer review ) in the ICDAR 2026 workshop "VINALDO: 3rd International Workshop on Machine Vision and NLP for Document Analysis"