超越盲目服从:OCR推理中的任务验证基准测试
Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning
浏览论文内容
中文总结 AI 辅助
该研究针对OCR推理中现有MLLMs的盲目服从问题,构建含1800样本的VeriOCRBench基准,评估15款领先MLLMs,发现其存在任务验证缺陷与过度弃权问题,揭示了OCR推理系统的可靠性差距。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在以OCR为核心的文档理解和文本丰富的视觉推理基准测试中已取得优异性能。然而现有评估大多假设每个任务都是有效且可回答的,而在现实世界的OCR场景中,这一假设往往不成立:问题可能依赖难以辨认的文本、被遮挡的证据、不存在的视觉目标、矛盾的前提或缺失的变量。我们将这一可靠性差距研究为基于OCR的任务验证:在回答前,模型应确定图像前提(IP)、文本前提(TP)和问题(Q)是否共同定义了一个可执行的任务。我们推出VeriOCRBench,这是一个包含1800个样本、经人工验证的基准测试,构建自8个OCR相关数据集的源图像,覆盖8个现实世界图像领域,并带有受控的、基于图像的诊断任务。它包含8种陷阱类型和4个验证维度——视觉、上下文、事实和逻辑——的1600个注入陷阱的无效任务,以及200个无陷阱的对照任务,用于衡量过度弃权(不执行)。该基准测试基于视觉原子事实(VAF)锚定的流程并经过全面人工审核,支持对任务验证、根本原因诊断和过度弃权进行解耦评估。对15个领先的MLLMs的评估显示出持续的盲目服从、诊断失败和提示诱导的过度弃权,暴露了当前OCR推理系统中的关键可靠性差距。代码可在以下网址获取:this https URL。
英文摘要
Multimodal Large Language Models (MLLMs) have achieved strong performance on OCR-centric document understanding and text-rich visual reasoning benchmarks. Yet existing evaluations largely assume that every task is valid and answerable. In real-world OCR scenarios, this assumption often fails: questions may rely on illegible text, occluded evidence, nonexistent visual targets, contradictory premises, or missing variables. We study this reliability gap as OCR-grounded Task Verification: before answering, a model should determine whether the Image Premise (IP), Textual Premise (TP), and Question (Q) jointly define an executable task. We introduce VeriOCRBench, a 1,800-sample human-verified benchmark built from source images drawn from 8 OCR-related datasets and spanning 8 real-world image domains, with controlled, image-grounded diagnostic tasks. It contains 1,600 trap-injected invalid tasks across 8 trap types and four verification dimensions---Visual, Contextual, Factual, and Logical---plus 200 trap-free controls for measuring over-refusal. Built with a Visual Atomic Fact (VAF)-anchored pipeline and full human auditing, VeriOCRBench enables decoupled evaluation of task verification, root-cause diagnosis, and over-refusal. Evaluating 15 leading MLLMs reveals persistent blind compliance, diagnosis failures, and prompt-induced over-refusal, exposing a critical reliability gap in current OCR reasoning systems. The code is available at: https://github.com/zy001122/Beyond-Blind-Compliance.
发表机构
- School of Artificial Intelligence, Jilin University(吉林大学人工智能学院)
- International Center of Future Science, Jilin University(吉林大学未来科学国际中心)
机构由 AI 辅助整理,请以论文原文为准。