HPMD:一个带行级标注的历史波斯语手稿数据集,用于词定位
HPMD: A Historical Persian Manuscript Dataset for Word Spotting with Line-Level Annotation
- Payame Noor University(帕亚梅努尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出历史波斯语手稿数据集HPMD(223页、3,678行),并设计仅用行级标注训练的词定位基线,通过后验矩阵匹配将F1从0.487提升至0.558,优于PHOC基线。
AI中文摘要:
大量历史波斯语手稿已被数字化,但检索它们仍然缓慢且大多依赖人工。历史学家通常希望找到特定人名、日期、事件或主题出现的位置,这是一个词定位问题。该任务的进展受限于两点:首先,几乎没有公开的历史波斯语手写数据集;唯一显著的资源OpenITI MAKHZAN仅包含相对较小的波斯语部分。其次,词定位模型通常需要词级边界框,而标注这些边界框非常昂贵。在本文中,我们引入了一个新数据集,包含223页、3,678行、37,631个词和130,630个字符,收集自多样的历史波斯语诗歌和散文书籍,并在区域、行和文本级别进行了标注。我们还提出了一个基线方法,该方法仅使用行级标注进行训练,但能返回词级位置。一个微调的行检测器用于查找文本行,一个微调的CRNN识别器(使用CTC训练)为每行生成逐帧字符后验矩阵。我们不解码每帧最可能的字符,而是直接针对该矩阵对查询进行评分,因此波斯语中视觉相似的字符(如be和pe)不再导致严重失败。帧对齐还提供了词在行内的水平位置。在测试集上,微调的行检测器达到F1为0.892,基于后验的搜索将词定位F1从0.487提高到0.558(与解码文本上的精确匹配相比),决策阈值在保留的验证集上选择。一个PHOC属性嵌入基线(在测试时额外接收oracle词边界)达到F1为0.449,低于所提方法。我们还报告了数据集的分布分析、检索错误分类以及按条件划分的性能细分。
英文摘要:
Large collections of historical Persian manuscripts have been digitized, but searching them is still slow and mostly manual. Historians usually want to find where a specific name, date, event, or topic appears, which is a word spotting problem. Progress on this task is limited by two things. First, there is almost no public dataset of historical Persian handwriting; the only notable resource, OpenITI MAKHZAN, contains a relatively small Persian portion. Second, word spotting models usually need word-level bounding boxes, which are very expensive to annotate. In this paper we introduce a new dataset of 223 pages, 3,678 lines, 37,631 words, and 130,630 characters, collected from diverse historical Persian books of poetry and prose and annotated at the region, line, and text level. We also propose a baseline that is trained only with line-level annotations but returns word-level locations. A fine-tuned line detector finds text lines, and a fine-tuned CRNN recognizer trained with CTC produces a frame-by-character posterior matrix for each line. Instead of decoding the most probable character at each frame, the query is scored directly against this matrix, so visually similar characters in Persian such as be and pe no longer cause hard failures. The frame alignment also gives the horizontal position of the word inside the line. On the test set, the fine-tuned line detector reaches an F1 of 0.892, and posterior-based search raises the word spotting F1 from 0.487 to 0.558 compared with exact matching on the decoded text, with the decision threshold selected on a held-out validation set. A PHOC attribute-embedding baseline that additionally receives oracle word boundaries at test time reaches an F1 of 0.449, below the proposed method. We also report a distributional analysis of the dataset, a taxonomy of retrieval errors, and a per-conditionbreakdown of performance.