哪种纸莎草手写文本识别(HTR)足够好?希腊文本四项纸莎草学任务对字符错误率的容忍度
Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts
浏览论文内容
中文总结 AI 辅助
本研究针对希腊纸莎草文本识别,测试不同字符错误率对四项纸莎草学任务的影响,发现各任务容忍度不同,且在含噪文本上重新训练模型可显著提升对不完美识别的鲁棒性。
中文摘要 AI 辅助
目的:大多数希腊纸莎草文献尚未出版和数字化;一种能够自动转录它们的笔迹文本识别(HTR)流程将使学者能够发现迄今未被阅读的文献和文学作品。古希腊纸莎草文献的识别系统尚处于起步阶段,对于给定的纸莎草学任务,其准确度要求尚未得到检验。为回答这一问题并为希腊纸莎草HTR设定基准,我们以已出版的版本为基准,测试了一系列字符错误率(CER)对四项纸莎草学任务的影响。方法:从this http URL中63,846个希腊文本的当前版本中,我们通过移除编辑层来模拟仅含字母的“完美HTR”输出,然后使用种子算法将其降级到精确的1%-50%的CER,包括丢失行和四种错误形态变体。在这些数据上,我们训练了用于文献类型、日期和文献与文学分类的小型模型(TF-IDF、fastText、字符CNN、ByT5-small),并应用了八种关键词搜索方法。我们将基于干净文本训练的模型与在特定CER水平下重新训练的模型进行比较,并在不同CER下进行评估。结果:不同任务的容忍度不同。对于基于干净文本训练的模型,文献与文学分类在高达20%的CER下仍保持其指标的90%;文献类型在高达7.5%的CER下保持;子类型和搜索在高达5%的CER下保持;日期仅在高达3%的CER下保持。在包含字符错误的文本上重新训练,在很大程度上消除了在超过15%的CER时出现的急剧退化。模型通常能更好地容忍长文档中的集中损伤,而不是短文本中分散的小错误。结论:该研究为四项任务中的每一项提供了CER目标,并表明在噪声文本上训练的模型使当前不完美的文本识别对这些任务变得有用。
英文摘要
Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would let scholars discover documents and literary works that have so far gone unread. Recognition systems for Ancient Greek papyri are in statu nascendi, and how accurate they must be for a given papyrological task has not been examined. To answer this and set a benchmark for Greek papyrus HTR, we test a range of character error rates (CER) against four papyrological tasks, using published editions as ground truth. Methods: From 63,846 current editions of Greek texts in papyri.info, we imitate a letters-only "perfect HTR" output by removing the editorial layer, then degrade it with a seeded algorithm to exact CERs of 1 - 50%, with lost lines and four error-shape variants. On these data we train small models (TF-IDF, fastText, a character CNN, ByT5-small) for document type, dating and documentary-versus-literary classification, and apply eight keyword search methods. We compare models trained on clean text with models retrained at a specific CER level, and evaluate across CERs. Results: Tolerance differs by task. With clean-trained models, documentary-versus-literary classification retains 90% of its metric up to 20% CER; document type up to 7.5%; subtypes and search up to 5%; dating only up to 3%. Retraining on text containing character errors largely eliminates the sharp degradation that otherwise sets in above 15% CER. Models generally tolerate concentrated damage in a long document better than small errors spread across a short text. Conclusion: The study provides a CER target for each of the four tasks and shows that models trained on noisy text make current, imperfect text recognition useful for them.
发表机构
- University of Vienna(维也纳大学)
机构由 AI 辅助整理,请以论文原文为准。