arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

史学资料数字化处理管道的对比评估

A Comparative Evaluation of Digitization Pipelines for Historiographical Sources

Marina Gómez Rey, Patricia Callejo, Mario Muñoz-Organero, Carlos Alario-Hoyos

arXiv 2608.24976首次发表:更新:

AI 中文总结

本研究对比评估了13种PDF转文本提取管道,发现Marker直接提取性能最优,LLM后校正无系统性改进,端到端解析是异构历史藏品的可靠方法。

AI 中文摘要

目的:历史文档的数字化为现代信息检索和人工智能(AI)系统带来了根本性挑战。源语料库中的光学字符识别(OCR)错误会通过检索增强生成(RAG)管道传播,损害生成输出的事实准确性。方法:本研究针对西哥特时期的史学二手资料,对PDF转文本提取管道进行了系统评估。我们评估了13种不同方法,涵盖三个方法族:直接提取、大语言模型(LLM)后校正,以及分块提取。根据制作方法和视觉复杂度,将文档分为五类。以人工校正的真值为基准,用字符错误率(CER)和单词错误率(WER)衡量性能。结果:结果表明,使用Marker进行直接提取表现出卓越性能(整体CER准确率达98.70%;WER准确率达97.71%),而传统OCR管道在扫描文档和复杂布局上表现出显著性能下降。嵌入文本提取在数字PDF上表现良好,但在扫描文档上失效。LLM后校正未提供系统性改进,且常降低提取准确性。结论:端到端文档解析是异构历史藏品最可靠的方法。扫描质量、布局复杂度和嵌入文本层的存在等文档特征对提取准确性有显著影响。基于LLM的后校正不应默认视为有益,大规模应用前需进行验证。

英文摘要

Purpose: The digitization of historical documents presents fundamental challenges for modern information retrieval and Artificial Intelligence (AI) systems. Optical character recognition (OCR) errors in source corpora propagate through retrieval-augmented generation (RAG) pipelines, compromising the factual accuracy of generated outputs. Methods: This study presents a systematic evaluation of PDF-to-text extraction pipelines applied to historiographical secondary sources on the Visigothic period. We assess thirteen distinct approaches spanning three methodological families: direct extraction, Large Language Model (LLM) post-correction, and chunk-and-extract. Documents are stratified into five categories based on production method and visual complexity. Performance is measured using character error rate (CER) and word error rate (WER) against manually corrected ground truth. Results: Results demonstrate that direct extraction with Marker achieves superior performance (98.70% CER accuracy; 97.71% WER accuracy overall), while conventional OCR pipelines exhibit substantial degradation on scanned documents and complex layouts. Embedded-text extraction performs well on digital PDFs but fails on scanned documents. LLM post-correction does not provide systematic improvements and frequently degrades accurate extractions. Conclusion: End-to-end document parsing is the most reliable approach for heterogeneous historical collections. Document characteristics such as scan quality, layout complexity, and the presence of embedded text layers have a significant impact on extraction accuracy. LLM-based post-correction should not be assumed beneficial by default and requires validation before large-scale application.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑