arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Institutional Books - Visual Elements:从数字藏书集中提取、分类、去重和生成视觉元素描述的开源管道

Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections

Jimmy Mendez, Matteo Cargnelutti, David Lowry-Duda, Catherine Brobston, Salwa Ismail, Greg Leppert, Amanda Watson, Jonathan Zittrain

arXiv 2608.18957首次发表:更新:

AI 中文总结

该研究推出开源管道Institutional Books - Visual Elements,可对历史藏书的视觉元素进行提取等处理,并发布含2260万个元素的数据集,助力数字化藏书的新计算应用。

AI 中文摘要

历史藏书包含丰富的视觉元素,如插图、照片、版画和装饰艺术,这些元素在大规模数字化项目中常未被充分利用。尽管光学字符识别(OCR)已实现文本内容提取的标准化,但这些视觉组件提供了自动化文本提取工作流程尚未挖掘的细微差别和上下文信息。本技术报告介绍了Institutional Books - Visual Elements,这是一个用于从历史藏书集中检测、分类、去重和生成视觉元素描述的开源端到端管道。伴随该管道,我们发布了一个初始数据集,包含从构成Institutional Books:哈佛图书馆数据集的983004册扫描卷中提取的2260万个视觉元素。本研究助力了社区范围内通过计算访问实现数字化藏书新用途的持续努力,涵盖从人工智能模型训练到数字人文研究等场景。

英文摘要

Historical book collections contain rich visual elements - such as illustrations, photographs, engravings, and decorative art - that are frequently under-explored in large-scale digitization projects. While Optical Character Recognition (OCR) has standardized the extraction of textual content, these visual components offer a layer of nuance and context that remains largely untapped by automated text extraction workflows. This technical report introduces Institutional Books - Visual Elements, an open-source end-to-end pipeline for detecting, classifying, deduplicating, and captioning visual elements from historical book collections. Alongside this pipeline, we release an initial dataset of 22.6 million visual elements extracted from the 983,004 scanned volumes that comprise the Institutional Books: Harvard Library dataset. This work contributes to ongoing, community-wide efforts to enable new use cases for digitized library collections through computational access, from artificial intelligence model training to digital humanities research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑