发表机构
Higher School of Computer Science (ESI-SBA)(高等计算机科学学院(ESI-SBA))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人员发布了最大规模带页边距和插入锚点标注的历史阿拉伯手稿行级数据集AraMS-28k,采用RefLAM流程构建,提供基线HTR结果,支持阿拉伯手稿相关研究。
AI 中文摘要
我们推出AraMS-28k,这是公开发布的最大规模真实历史阿拉伯手稿行级数据集,包含14部书籍、3043页内容及28600条标注文本行(其中27971条为主文本行,629条为页边距行)。13部为手写手稿,涵盖Naskh、Ruq'ah和Maghrebi三种字体传统,1部为石印印刷版本以扩大格式多样性。每条行标注为主文本或页边距,且页边距行若在主文本中有明确附着点,还会标注插入锚点,在行级粒度上恢复手稿真实的非线性阅读顺序——据我们所知,这是首个针对历史阿拉伯手稿语料库发布的此类标注。由于参考转录完全标音,而手稿笔迹通常未标音符,我们为每条行同时发布原始标音转录和标音归一化版本。该数据集通过RefLAM构建,这是一种基于参考的标注流程,将多模态大语言模型OCR与独立来源的干净转录对齐,并对每条行进行人工审核,结合自动验证与专家监督。我们描述了构建和质量控制流程,呈现了标注模式,报告了语料库和单书籍层面的数据集统计数据,并提供了使用Kraken和HATFormer的基线HTR结果,包括从分布内页面到完全未见书籍的跨字体泛化梯度。AraMS-28k以页面图像、行级标注和固定的训练/验证/测试集拆分形式发布,采用CC BY-NC-SA 4.0许可,以支持阿拉伯手稿识别、布局分析和阅读顺序恢复的可复现研究。
英文摘要
We introduce AraMS-28k, the largest publicly released line-level dataset of genuine historical Arabic manuscripts, comprising 14 books, 3,043 pages, and 28,600 annotated text lines (27,971 main-text, 629 margin). Thirteen books are hand-copied manuscripts spanning three script traditions -- Naskh, Ruq'ah, and Maghrebi -- and one is a lithographed printed edition included to broaden format diversity. Each line is labelled as main-text or margin, and margin lines that have an unambiguous attachment point in the main text are further annotated with an insertion anchor, recovering the manuscript's true non-linear reading order at line-level granularity -- to our knowledge the first such annotation released for a historical Arabic manuscript corpus. Because reference transcriptions are fully vocalised while manuscript hands are typically undiacritised, we release both the raw diacritised transcription and a diacritic-normalised counterpart for every line. The dataset was constructed with RefLAM, a reference-grounded annotation pipeline that aligns multimodal-LLM OCR against independently sourced clean transcriptions and routes every line through human review, combining automatic verification with expert oversight. We describe the construction and quality-control process, present the annotation schema, report dataset statistics at both the corpus and per-book level, and provide baseline HTR results using Kraken and HATFormer, including a cross-script generalisation gradient from in-distribution pages to fully unseen books. AraMS-28k is released with page images, line-level annotations, and fixed train/val/test splits under CC BY-NC-SA 4.0 to support reproducible research on Arabic manuscript recognition, layout analysis, and reading-order recovery.
CommentsData and code available at this https URL and this https URL. Dataset: https://doi.org/10.5281/zenodo.22095333 Code: https://github.com/ArchaText/AraMS-28k-Dataset