arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03617cs.CLcs.CVcs.DL

俄罗斯科学院档案馆齐奥尔科夫斯基论文(档案全宗555)的机器可读目录及其手写文本可读性测量方法

A machine-readable catalogue of the Tsiolkovsky papers (fond 555, Archive of the Russian Academy of Sciences), and a way to measure how well its handwriting can be read

Vladimir Beskorovainyi

首次发表
浏览论文内容

中文总结 AI 辅助

本文构建了齐奥尔科夫斯基论文的机器可读目录,提出了无真实标注下的手写文本识别准确率测量方法,并验证了其应用边界。

中文摘要 AI 辅助

康斯坦丁·齐奥尔科夫斯基(1857-1935)的个人档案作为全宗555保存于俄罗斯科学院档案馆。该档案馆已对该全宗进行扫描并发布了图像,但无查询目录、无全文搜索、无数据集,仅能逐页浏览馆藏。本文描述了一个包含全部2019份文件、51008幅扫描件的机器可读目录,其中1969份文件的日期取自档案馆自身描述,每幅扫描件均按手写体和印刷体进行了页级分类,还有一个不断扩充的机器转录语料库(目前包含322份文件、5454幅扫描件)。本文还提出了一种在无真实标注的档案馆中测量手写文本识别准确率的方法。打字机时代的档案馆常将同一文本以手稿和打字副本两种形式保存,对两者进行转录并对比可分离出阅读误差,因为来源和处理流程完全相同,仅页面难度存在差异。在来自27份文件的294对此类样本中,两份手写页面的读数在中位数37%的单词上一致。对于两份还有出版版本的文件,可对照真实标注验证该估计值:其偏差在1个百分点以内,且对页面的排序与真实标注一致(当出版版本为可靠见证时,秩相关系数为0.92)。这限定了其应用:本文中某作品的两个版本共享19%的单词,低于单页两次读数的一致率,因此无法以该质量逐字核对修订本。该负面结果已如实报告,且该约束已嵌入工具中。

英文摘要

The personal archive of Konstantin Tsiolkovsky (1857-1935) is held as fond 555 of the Archive of the Russian Academy of Sciences. The archive scanned the fond and published the images, but with no queryable catalogue, no full-text search and no dataset: the holdings can only be browsed one page at a time. This paper describes a machine-readable catalogue of all 2,019 files and 51,008 scans, a dating for 1,969 files taken from the archive's own descriptions, a page-level classification of every scan into handwriting and typescript, and a machine transcription of the fond in full: all 2,019 files and all 51,008 scans. It also reports a way to measure handwritten-text-recognition accuracy in an archive with no ground truth. Archives of the typewriter era often preserve one text twice, as manuscript and as a typed copy; transcribing both and comparing isolates the reading error, since source and pipeline are identical and only page difficulty differs. Across 1,759 such pairs from 224 files, two readings of a handwritten page agree on a median 37% of words and share a longest verbatim run of a median 10 words; the median is unchanged from the 294 pairs of the first version, on a sample six times larger. On two files that also have a published edition the estimate can be checked against ground truth: it is unbiased to within a percentage point and ranks pages as the truth does (rank correlation 0.92 where the edition is a faithful witness). This bounds use: two variants of one work here share 19% of words, below the rate at which two readings of a single page agree, so the redactions cannot be collated word by word at this quality. That negative result is reported as such, and the constraint is built into the tool.

补充信息

↑