当视觉语言模型信任上下文:评估误导性上下文下的场景文本识别
When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context
浏览论文内容
中文总结 AI 辅助
本研究提出SceneFaith基准,揭示视觉语言模型在场景文本识别中受上下文误导而改写文本,并强调需平衡视觉证据与上下文信息。
中文摘要 AI 辅助
视觉语言模型(VLMs)能够阅读自然场景中的文本,但其预测可能受到周围上下文的影响。当印刷文本与场景所暗示的内容相冲突时,模型可能会返回一个更合理的词,而非实际显示的文本。我们引入了SceneFaith,一个包含781张生成场景图像的基准,用于研究这一行为。每个输出被分类为字面(Literal)、规范(Canonical)或其他(Other),以区分忠实转录、上下文一致的改写和普通识别错误。在来自七个家族的15个模型中,所有模型在清晰图像上均表现出改写行为,改写率介于8.45%至58.51%之间。受控实验进一步表明,周围上下文至关重要:移除周围场景信息会减少改写并提高字面准确率,而改变同一文本块周围的场景也会改变模型输出。此外,通过模糊处理削弱目标文本会增加改写。这些结果表明,可靠的场景文本识别要求视觉语言模型在视觉字符证据与上下文信息之间取得平衡,在保留清晰文本的同时,主要仅在视觉证据不确定时使用上下文。
英文摘要
Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45\% to 58.51\%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.
发表机构
- Jilin University(吉林大学)
机构由 AI 辅助整理,请以论文原文为准。