arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24077cs.CVcs.LGeess.IV

当低字符错误率还不够时:对乌拉圭历史文档视觉语言OCR系统中幻觉的分析

When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

Marina Gardella, Camilo Mariño, Diego Belzarena, Ignacio Ramírez, Gregory Randall, Jean-Michel Morel

首次发表
浏览论文内容

中文总结 AI 辅助

研究乌拉圭历史文档视觉语言OCR系统,在Berrutti数据集上对比传统OCR与VLM方法,发现VLM虽CER和WER优,但存在系统性失败模式,揭示了定量性能与转录保真度差距,强调需超越字符级准确性的评估框架。

中文摘要 AI 辅助

光学字符识别(OCR)是历史档案数字化的关键组成部分。近来,视觉语言模型(VLM)成为传统OCR系统的有力替代品,在标准基准测试中取得了领先性能。然而,其在档案转录方面的适用性仍未得到充分理解。本文在Berrutti数据集(一个源自缩微胶卷扫描的具有挑战性的乌拉圭独裁统治时期文档集合)上对传统OCR系统和基于VLM的方法进行了基准测试。虽然VLM在字符错误率(CER)和单词错误率(WER)方面始终优于传统方法,但这些改进掩盖了更复杂的情况。通过详细的定性分析,我们发现了标准度量无法察觉的系统性失败模式。影响命名实体的错误尤为关键,因为它们会在对CER和WER影响极小的情况下引入大量语义扭曲。这些发现揭示了现实档案环境中定量OCR性能与转录保真度之间的关键差距,并强调需要超越字符级准确性的评估框架来捕捉生成转录的语义可靠性。

英文摘要

Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on standard benchmarks. However, their suitability for archival transcription remains insufficiently understood. In this work, we benchmark traditional OCR systems and VLM-based approaches on the Berrutti dataset, a challenging collection of Uruguayan dictatorship-era documents derived from microfilm scans. While VLMs consistently outperform traditional methods in terms of Character Error Rate (CER) and Word Error Rate (WER), we show that these improvements hide a more complex picture. Through a detailed qualitative analysis, we uncover systematic failure modes that are invisible to standard metrics, including orthographic normalization, spurious content generation, and semantic substitutions that preserve fluency while altering meaning. Errors affecting named entities are particularly critical, as they can introduce substantial semantic distortions with minimal impact on CER and WER. These findings reveal a critical gap between quantitative OCR performance and transcription fidelity in real-world archival settings, and highlight the need for evaluation frameworks that go beyond character-level accuracy to capture the semantic reliability of generated transcriptions.

发表机构

  • Université Paris-Saclay(巴黎-萨克雷大学)
  • ENS Paris-Saclay(巴黎-萨克雷高等师范学校)
  • CNRS(法国国家科学研究中心)
  • Centre Borelli(博雷利中心)
  • Facultad de Ingeniería, Universidad de la República(乌拉圭共和国大学工程学院)
  • Lingnan University(岭南大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑