AI 中文总结
研究基于视觉语言模型的端到端OCR系统效率问题,提出LayoutLite模块,通过令牌级隐式布局分析,聚合视觉表示并预测重要性分数,去除低信息令牌,实验证明其能有效加速OCR系统,提升效率且保证识别质量。
AI 中文摘要
基于视觉语言模型的端到端OCR系统在复杂文档OCR中表现出色,但效率受文档图像产生的大量视觉令牌限制。许多令牌对应空白边缘或视觉冗余区域,直接应用通用视觉令牌压缩方法可能去除关键OCR细粒度细节。本文提出LayoutLite,一个用于高效文档OCR的轻量级即插即用模块。它在视觉编码器和语言解码器之间的令牌级别执行隐式布局分析,聚合视觉编码器的多层视觉表示,用轻量级评分网络预测每个视觉令牌的重要性分数,在进入语言解码器前去除低信息令牌并保留其空间位置信息。为在无人工标注下训练LayoutLite,将令牌选择视为强化学习问题,用由OCR输出一致性驱动的组相对策略优化目标及辅助布局监督信号进行优化。在OmniDocBench上的实验表明LayoutLite能大幅减少视觉令牌长度和推理成本,识别质量下降可忽略不计。在两个OCR专用VLM上进一步评估,在高达50%的令牌压缩下,LayoutLite在两个模型上保持几乎相同分数,同时将预填充延迟、FLOPs和KV缓存内存减少超40%,仅增加少量推理开销。这些结果表明令牌级隐式布局分析是加速基于VLM的OCR系统的有效实用方法。
英文摘要
End-to-end OCR systems based on vision-language models have achieved strong performance in complex document OCR, but their efficiency is limited by the large number of visual tokens produced from document images. Many of these tokens correspond to blank margins or visually redundant regions, yet directly applying generic visual token compression methods may remove OCR-critical fine-grained details. In this paper, we propose LayoutLite, a lightweight plug-and-play module for efficient document OCR. Instead of relying on explicit document layout detection, LayoutLite performs implicit layout analysis at the token level between the vision encoder and the language decoder. It aggregates multi-layer visual representations from the vision encoder, and predicts an importance score for each visual token with a lightweight scoring network. Low-information tokens are then removed before entering the language decoder while preserving the original spatial positional information of retained tokens. To train LayoutLite without human annotations, we cast token selection as a reinforcement learning problem and optimize it with a group-relative policy optimization objective driven by OCR output consistency, together with an auxiliary layout supervision signal to stabilize training. Experiments on OmniDocBench demonstrate that LayoutLite can substantially reduce visual token length and inference cost with negligible degradation in recognition quality. We further evaluate LayoutLite on two OCR-specialized VLMs, FireRed-OCR and Logics-Parsing-V2. Under up to 50% token compression, LayoutLite preserves almost the same score on both models while reducing prefill latency, FLOPs, and KV cache memory by over 40%, with only a small additional inference overhead. These results show that token-level implicit layout analysis is an effective and practical approach for accelerating VLM-based OCR systems.
Comments11 pages, 7 figures