arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从检索到识别:视觉-语言模型如何成为OCR专家

From Retrieval to Recognition:How Vision--Language Models Become OCR Specialists

Yuanxiang Huangfu, Hanmeng Zhong, Linqing Chen, Jeffrey Tiong Jee Hui

arXiv 2609.21543首次发表:更新:

发表机构

PatSnap Co., LTD.(智慧芽有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过因果干预和头部分析,发现通用视觉-语言模型通过重用文本检索/复制头来实现全序列OCR,并在专门化时保留头身份但重新分配其功能强度。

AI 中文摘要

一个通用的视觉-语言模型是通过发展新的阅读回路,还是通过重用现有机制来获得专门的OCR能力?我们在全序列OCR(而非局部答案检索)的背景下解决这个问题。使用基于证据的协议和保留的因果干预,我们在GLM-OCR、MinerU2.5和PaddleOCR-VL-1.6中识别出稀疏且稳定的OCR头集合。然后,我们通过将这些OCR头与在通用VLM中独立识别的文本检索/复制头进行比较,来研究这些OCR头的机制起源。在两个通用VLM中,视觉OCR头与独立识别的文本检索/复制头强烈重叠,未调优的前20个交集为73.3%,所有头的Spearman相关系数为0.677-0.886。重叠和因果干预表明,全序列OCR作为密集的序列多模态复制粘贴操作,反复检索视觉证据并将其路由到当前输出位置。最后,我们研究了当通用VLM成为OCR专家时,这个共享回路如何变化。匹配的基础到专门化的比较表明,OCR专门化在很大程度上保留了头的身份,每个任务保留前20个头中的17-20个,所有头的秩相关系数为0.874-0.942,同时重新分配了它们的功能和因果强度。

英文摘要

Does a general vision--language model acquire specialized OCR ability by developing a new reading circuit or by reusing an existing mechanism? We address this question in the setting of full-sequence OCR, rather than local-answer retrieval. Using an evidence-grounded protocol with held-out causal interventions, we identify sparse and stable OCR-head sets in GLM-OCR, MinerU2.5, and PaddleOCR-VL-1.6. We then investigate the mechanistic origin of these OCR heads by comparing them with independently identified textual retrieval/copy heads in general VLMs. Across two general VLMs, visual OCR heads strongly overlap independently identified textual retrieval/copy heads, yielding untuned top-20 intersections of 73.3% and all-head Spearman correlations of 0.677-0.886. The overlap and causal interventions suggest that full-sequence OCR operates as dense sequential multimodal copy-and-paste, repeatedly retrieving visual evidence and routing it to the current output position. Finally, we examine how this shared circuit changes as a general VLM becomes an OCR specialist. Matched base-to-specialized comparisons show that OCR specialization largely preserves head identity, retaining 17-20 of the top 20 heads per task with all-head rank correlations of 0.874-0.942, while redistributing their functional and causal strengths.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑