arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19026cs.CLcs.DL

机构藏书-丰富文本:一个可定制的多语言开源流水线,用于大规模对OCR文本进行去噪、去重和标注

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

David Lowry-Duda, Matteo Cargnelutti, Catherine Brobston, Salwa Ismail, Greg Leppert, Amanda Watson, Jonathan Zittrain

AI总结:

该研究提出名为Enriched Text的多语言开源流水线,可大规模对IB-HL的OCR文本去噪、去重、标注,保留元数据,支持用户定制输出,发布了IB-HL-ET版本以提升藏书集的机器可读性与研究价值。

AI中文摘要:

2025年发布的哈佛图书馆机构藏书(IB-HL)是一个包含983004卷的藏书集(共242B o200k_base tokens),最初通过哈佛图书馆参与谷歌图书馆项目完成数字化。随着研究人员和开发者开始使用IB-HL,大规模预处理的标准实践与精细信息管理目标之间出现了矛盾。许多现有流水线针对网页文本进行优化,因此倾向于激进地过滤、去重、按语言限制,有时还会丢弃有价值的元数据;与此同时,试图使用IB-HL的研究人员在进行类似处理和分析时重复付出精力。我们描述了一种称为丰富文本(Enriched Text)的方法:该方法不生成单一“完整”的token流,而是在保留元数据的同时对文本进行归一化处理,具体操作包括分离末尾内容、检测段落级语言、识别重复段落簇、计算段落级每字节比特数得分,并通过叠加在文本上的类HTML标注提供这些信息;用户可通过解析这些标注,根据自身需求定制输出,而非接受对内容的全局编辑决策。该流水线适用于藏书集中约250种语言。本报告阐述了该项目的目标、实现方式和设计原理,发布内容包括IB-HL-ET(IB-HL的丰富文本版本,包含983003卷共217B o200k_base tokens,组织为13.9亿个带标注的子主题段落)以及生成该版本的流水线,这些内容旨在让该藏书集更便于机器解析和人类研究。

英文摘要:

Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As researchers and developers have begun to use IB-HL, a tension has emerged between standard large-scale preprocessing practices and the goals of careful information stewardship. Many existing pipelines optimize for web text: as a result, they tend to aggressively filter, deduplicate, restrict by language, and sometimes discard meaningful metadata. Meanwhile, researchers seeking to use IB-HL duplicate effort while performing similar processing and analysis. We describe an approach that we call Enriched Text. Instead of producing a single 'complete' stream of tokens, we normalize the text while preserving metadata through annotations. We separate endmatter, detect per-paragraph language, identify clusters of duplicate paragraphs, and compute per-paragraph bits-per-byte scores. We provide this information through HTML-like annotations layered on top of the text. By parsing these annotations, users can tailor the output to their own needs instead of accepting a global editorial decision on content. The pipeline applies to all $\approx$250 languages in the collection. This report describes this project's goals, implementation, and design rationale. The release includes IB-HL-ET (an enriched-text version of IB-HL containing 217B o200k_base tokens across 983,003 volumes, organized into 1.39B annotated subtopic paragraphs) and the pipeline that produced it. These serve to make the collection easier for machines to parse and for humans to study.

↑