AI 中文总结
研究视觉语言模型在文档理解中是否忠实转录。引入FaithC4基准,评估通用、OCR专用VLMs及传统OCR管道。发现通用VLMs受扰WER降4.5分,OCR专用降0.2 - 2分,传统OCR降不到0.6分。还揭示重写与FFN表示及单词长度的关系。
AI 中文摘要
视觉语言模型(VLMs)越来越多地用于取代传统的光学字符识别(OCR)管道进行文档理解。本文表明它们并不总是忠实的转录器:当文本不完美时,它们往往会将其重写成更合理的形式,这种行为是干净文本OCR基准无法检测到的。我们引入了FaithC4,这是一个包含1455个单页文档(英语、中文、韩语)的多语言扰动基准,有三种扰动类型:加扰、随机替换和视觉相似替换。我们用该基准评估了15个系统,包括通用VLMs、OCR专用VLMs和传统OCR管道。这三类在扰动下的字错误率(WER)下降情况不同:通用VLMs下降高达4.5个百分点,OCR专用VLMs下降0.2 - 2个百分点,传统OCR在英语上下降不到0.6个百分点。逐层探测Qwen3 - VL - 4B,我们发现了一个一致的模式:只有当扰动单词的最后一层前馈神经网络(FFN)表示与原始编码接近时,重写才会发生;当表示差异足够大时,模型会忠实地转录。单词长度影响重写率:短单词(4 - 6个字符)重写率高达10%,8个字符以上急剧下降至0%。
英文摘要
Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding. In this paper, we show they do not always act as faithful transcribers: when text is imperfect, they often tend to rewrite it into a more plausible form - a behavior that clean-text OCR benchmarks cannot detect. We introduce FaithC4, a multilingual perturbation benchmark of 1,455 single-page documents (English, Chinese, Korean) with three perturbation families: scramble, random substitution, and visually similar substitution. We use the benchmark to evaluate 15 systems spanning general-purpose VLMs, OCR-specialized VLMs, and traditional OCR pipelines. These three categories differ in WER degradation under perturbation: general-purpose VLMs degrade by up to 6.9 points, OCR-specialized VLMs by 0.1-3.4 points, and traditional OCR by less than 0.8 points on English. Probing Qwen3-VL-4B layer-by-layer, we identify a consistent pattern: rewriting fires only when a perturbed word's final layer FFN representation stays close to the original encoding; when the representation diverges sufficiently, the model transcribes faithfully. Word length affects rewriting rate: short words (4-6 characters) are rewritten up to 10% of the time, with a sharp cutoff at 8 characters above which rewriting drops to 0%.
Comments15 pages, 6 figures