WildHandBench:挑战多模态大语言模型(MLLMs)与人类的手写文本理解基准
WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans
浏览论文内容
中文总结 AI 辅助
本文提出WildHandBench手写文本理解基准,引入PDE指标评估模型错误成因,实验发现最优MLLMs仅达71.85%准确率,人类表现更优但差距小,模型错误多由语言先验导致。
中文摘要 AI 辅助
尽管OmniDocBench上的顶尖模型在印刷文档解析任务中已达到96.34%的整体准确率,但当前模型处理具有挑战性的手写文档的能力仍未得到充分表征。现有基准多聚焦于孤立文本或公式,忽略手写表格与真实场景退化问题,且仅报告整体准确率而未说明模型失败的原因。本文提出WildHandBench,这一基准包含500份手写文档,涵盖自由文本、表格、公式三种结构,四种语言,以及九种真实场景。我们引入先验驱动误差(Prior-Driven Error,PDE)指标,用于量化模型错误是否源于语言先验而非视觉证据。通过评估18个最先进模型及校准后的人类基线,我们发现:(1)最优模型仅达到71.85%的整体准确率;(2)人类表现优于所有模型,但差距较小(77.09%对71.85%);(3)模型错误与人类错误存在质的差异——63%-91%的模型错误为先验驱动,而人类仅为49%,这暴露了模型对语言先验的系统性依赖,而传统准确率指标无法捕捉这一问题。
英文摘要
While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten documents across three structures (free text, tables, formulas), four languages, and nine real-world scenarios. We introduce a Prior-Driven Error (PDE) metric that quantifies whether errors originate from language priors rather than visual evidence. Evaluating 18 state-of-the-art models together with calibrated human baselines, we find: (1) the best model achieves only 71.85% overall; (2) humans outperform all models yet the gap is narrow (77.09% vs. 71.85%); and (3) model errors are qualitatively different from human errors -- 63-91% of model errors are prior-driven versus only 49% for humans, exposing systematic reliance on language priors that conventional accuracy metrics cannot capture.