发表机构
Institute of Automation, CAS; School of Artificial Intelligence, University of Chinese Academy of Sciences; Mohamed bin Zayed University of Artificial Intelligence; Zhongke Fanyu Technology Co., Ltd(中国科学院自动化研究所; 中国科学院大学人工智能学院; 穆罕默德·本·扎耶德人工智能大学; 中科梵语科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出ReViCo基准,以视觉文本纠错任务评估VLMs,发现其与人类存在显著性能差距,为开发更优文本感知VLMs提供新基础。
AI 中文摘要
视觉语言模型(Vision Language Models,VLMs)在通用视觉任务中已展现出出色性能,但仍难以深入理解图像中的文本。本文提出ReViCo(Real Visual Correction,真实视觉纠错),这是一个通过视觉文本纠错任务评估VLM文本理解能力的基准。ReViCo要求模型识别并修正真实图像中的文本错误,需深刻理解视觉文本与周围视觉上下文的关联。我们采用提示策略和针对性模型训练两种不同范式对多种VLM进行基准测试,以挖掘当前模型的极限。实验结果显示,即使是最优VLM与人类之间也存在显著性能差距;进一步分析表明,多数模型难以准确感知视觉文本,导致频繁出现修正错误。ReViCo通过凸显这些差距,为开发更鲁棒、更具文本感知能力的VLM提供了新的基准基础。
英文摘要
Vision Language Models (VLMs) have shown great success in general visual tasks, yet they still struggle to deeply understand text within images. In this paper, we introduce ReViCo (Real Visual Correction), a benchmark designed to evaluate VLM text understanding through a novel task of visual text error correction. ReViCo challenges models to identify and fix text errors in real-world images, which requires a profound understanding of the interplay between visual text and its surrounding visual context. We benchmark various VLMs using two distinct paradigms: prompt-based strategy and targeted model training, both aimed at pushing the limits of current models. Our experiments reveal a striking performance gap between even the best VLMs and human, and further analysis also shows that most models struggle to accurately perceive the visual text, resulting in frequent correction errors. By highlighting these gaps, ReViCo provides a new benchmark foundation for developing more robust and text-aware VLMs.