arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NaviDC-OCR:面向数字文档与相机拍摄文档的解析导航

TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, MingKun Jiang, Zhongjiang He, Hao Sun

arXiv 2608.12898首次发表:更新:

发表机构

China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd.(中国电信人工智能技术(北京)有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有文档解析方法的两大挑战,本文提出NaviDC-OCR框架,通过形变感知学习、自适应采样及内容-结构解耦学习策略,在多基准测试中取得最优性能并获ICDAR 2026相关挑战第一。

AI 中文摘要

文档解析旨在将非结构化文档转换为结构化、可被机器读取的表示形式。视觉语言模型(VLM)的近期进展大幅推动了文档解析技术,但现有方法仍面临两大核心挑战:其一,基于解耦VLM的方法高度依赖准确的布局分析,而相机拍摄文档中的几何畸变会引发级联误差;其二,尽管端到端VLM方法降低了对显式布局检测的依赖,但在高分辨率场景中常出现冗余生成、幻觉及结构推理不足的问题。为应对上述挑战,本文提出统一文档解析框架NaviDC-OCR,该框架引入形变感知学习以将几何感知融入VLM,并针对复杂布局表示提出自适应采样机制;此外,开发了内容-结构解耦学习策略,以显式建模公式语法与表格结构,实现更高效的结构化表示学习。大量实验表明,NaviDC-OCR在各类文档解析基准上均取得最优性能:在OmniDocBench v1.6、Wild-OmniDocBench、PureDocBench上分别获得96.87、88.53、78.41的整体得分,且在ICDAR 2026 Sci-ImageMiner挑战赛中排名第一,验证了NaviDC-OCR在复杂文档解析场景中的有效性与泛化能力。

英文摘要

Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose TeleOCR, a unified framework for document parsing. TeleOCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that TeleOCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of TeleOCR in complex document parsing scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑