机构报纸流水线:从历史报纸中提取数十亿高质量 token
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
浏览论文内容
中文总结 AI 辅助
该研究提出了与波士顿公共图书馆联合设计的Institutional Newspapers Pipeline模块化系统,可从历史报纸扫描件中提取高质量结构化数据集,已产出含163亿o200k_base token的开放数据集,为解锁海量报纸高质量数据迈出重要一步。
中文摘要 AI 辅助
历史报纸是公共生活的丰富记录,但其密集、不规则且有时嘈杂的布局使得对这些材料的计算访问既具有挑战性又受到限制。我们提出了机构报纸流水线(Institutional Newspapers Pipeline),这是我们与波士顿公共图书馆(Boston Public Library)联合设计的模块化系统,用于从历史报纸扫描件中提取高质量、结构化的数据集。该流水线的架构确保每个步骤都保持可解释性和可定制性,并且整个流水线在工作站级硬件上运行时仍保持计算经济性。流水线对每个扫描件执行多步骤流程:将扫描件分割为与类型无关的单个裁剪区域,对每个生成的裁剪区域执行光学字符识别(OCR),之后对每个裁剪区域执行文本分析、类型分类、阅读顺序检测、命名实体识别、主题分类、语言检测以及预计算嵌入生成。我们将该流水线应用于波士顿公共图书馆藏品的一部分,并将结果作为开放数据集发布。光学字符识别(OCR)输出代表了1795年至1930年发布的1,473,635份公共领域报纸扫描件中提取的8310万个独立裁剪区域内的163亿个o200k_base token。本报告描述了我们为每个处理步骤采用的方法、我们训练的小型模型,以及我们在此过程中收集的评估结果和数据集规模测量数据。它随流水线、模型和数据集的发布一同提供。我们认为这项工作是解锁数千万份报纸扫描件中高质量数据的重要一步。
英文摘要
Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts make computational access to these materials both challenging and limited. We present the Institutional Newspapers Pipeline, a modular system we jointly designed with Boston Public Library to extract high-quality, structured datasets from historical newspaper scans. It was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware. The pipeline runs each scan through a multi-step process: it segments scans into individual type-agnostic crops and performs OCR on each resulting segment before then performing text analysis, type classification, reading order detection, named entities recognition, subject classification, language detection, and pre-computed embeddings generation on every crop. We ran this pipeline against a portion of Boston Public Library's holdings and released the results as an open dataset. The optical character recognition (OCR) output represents 16.3 billion o200k_base tokens across 83.1 million individual crops, extracted from 1,473,635 public domain newspaper scans published between 1795 and 1930. This report describes our methods for each processing step, the small models we trained, as well as the evaluation results and dataset-scale measurements we collected in the process. It accompanies the release of the pipeline, models, and dataset. We position this work as a substantial step towards unlocking high-quality data from tens of millions of newspaper scans.