arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22959cs.CVcs.AI

WildHandBench:挑战多模态大语言模型(MLLMs)与人类的手写文本理解基准

WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

Jun Zhang, Qiao Zhao, Cheng Cui, Jianying Qu, Zhongkai Sun, Jianwen Yang, Changda Zhou, ZhuoXin Liu, Shubin Han

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出WildHandBench手写文本理解基准,引入PDE指标评估模型错误成因,实验发现最优MLLMs仅达71.85%准确率,人类表现更优但差距小,模型错误多由语言先验导致。

中文摘要 AI 辅助

尽管OmniDocBench上的顶尖模型在印刷文档解析任务中已达到96.34%的整体准确率,但当前模型处理具有挑战性的手写文档的能力仍未得到充分表征。现有基准多聚焦于孤立文本或公式,忽略手写表格与真实场景退化问题,且仅报告整体准确率而未说明模型失败的原因。本文提出WildHandBench,这一基准包含500份手写文档,涵盖自由文本、表格、公式三种结构,四种语言,以及九种真实场景。我们引入先验驱动误差(Prior-Driven Error,PDE)指标,用于量化模型错误是否源于语言先验而非视觉证据。通过评估18个最先进模型及校准后的人类基线,我们发现:(1)最优模型仅达到71.85%的整体准确率;(2)人类表现优于所有模型,但差距较小(77.09%对71.85%);(3)模型错误与人类错误存在质的差异——63%-91%的模型错误为先验驱动,而人类仅为49%,这暴露了模型对语言先验的系统性依赖,而传统准确率指标无法捕捉这一问题。

英文摘要

While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten documents across three structures (free text, tables, formulas), four languages, and nine real-world scenarios. We introduce a Prior-Driven Error (PDE) metric that quantifies whether errors originate from language priors rather than visual evidence. Evaluating 18 state-of-the-art models together with calibrated human baselines, we find: (1) the best model achieves only 71.85% overall; (2) humans outperform all models yet the gap is narrow (77.09% vs. 71.85%); and (3) model errors are qualitatively different from human errors -- 63-91% of model errors are prior-driven versus only 49% for humans, exposing systematic reliance on language priors that conventional accuracy metrics cannot capture.

↑