arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07917cs.AI

TongGuOCR:面向中文历史文献的布局感知与 token 增强 OCR 框架

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin

首次发表
浏览论文内容

中文总结 AI 辅助

针对中文历史文献OCR的复杂布局、生僻字等挑战,提出TongGuOCR框架,通过布局感知预处理与token增强识别模块,在M5HisDoc等基准上取得优于同类模型的性能。

中文摘要 AI 辅助

中文历史文献承载着珍贵的文化遗产,但许多馆藏仅以扫描页图像形式存在,无法实现全文检索、校对及计算分析。光学字符识别(OCR)可弥合这一差距,但因历史文献常包含复杂布局、生僻字符及不规则阅读顺序,准确转录仍具挑战性。我们提出 TongGuOCR——面向中文历史文献的布局感知与 token 增强 OCR 框架。首先,布局感知预处理模块构建并优化局部连贯识别块,以保留局部上下文同时减少区域间干扰。其次,token 增强识别模块在两个互补层面增强转录目标:字符级词汇扩展为每个生僻字提供直接的单 token 表示并缩短其解码路径;行间过渡建模注入离散空间位移 token,引导解码器沿复杂阅读路径,无需精确坐标。在两个中文历史文献 OCR 基准上的实验表明,TongGuOCR 优于代表性的传统专用 OCR 模型、通用多模态大语言模型(MLLMs)及面向 OCR 的 MLLMs。在更具挑战性的 M5HisDoc 基准上,TongGuOCR 达到 93.76 的 AR,且相对于各指标的最佳竞争得分,将 NED 从 10.43 降至 6.15,RO-ED 从 7.53 降至 3.49。在线演示可通过此 URL 获取。

英文摘要

Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis. Optical character recognition (OCR) can bridge this gap, but accurate transcription remains challenging because historical documents often contain complex layouts, rare characters, and nontrivial reading orders. We propose TongGuOCR, a layout-aware and token-augmented multimodal large language model (MLLM) for OCR of Chinese historical documents. First, a Layout-Aware Preprocessing module constructs and refines locally coherent recognition blocks to preserve local context while reducing interference across regions. Second, a Token-Augmented Recognition module augments the transcription target at two complementary levels: character-level vocabulary expansion gives each rare glyph a direct one-token representation and shortens its decoding path, while line-to-line transition modeling injects discrete spatial displacement tokens that guide the decoder along complex reading paths without requiring precise coordinates. Experiments on two Chinese historical document OCR benchmarks show that TongGuOCR outperforms representative traditional task-specific OCR models, general-purpose MLLMs, and OCR-oriented MLLMs. On the more challenging M5HisDoc benchmark, TongGuOCR achieves 93.76 AR and reduces NED from 10.43 to 6.15 and RO-ED from 7.53 to 3.49 relative to the best competing score for each metric. An online demo is available at https://jzzh2004.github.io/TongGuOCR.

发表机构

  • School of Electronic and Information Engineering, South China University of Technology(华南理工大学电子与信息工程学院)
  • Huawei Technologies Co., Ltd.(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑