arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07943cs.AI

多页视觉丰富文档理解中的故障定位:实证归因

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

Lewei Xu, Yihao Ding, Zihan Xu, Daniel Yitian Su, Daochang Liu, Siwen Luo, Yifan Peng, Wei Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对MP-VRDU的错误答案,分离出表示、选择、推理三种故障模式,发现视觉信息不可替代、缺失页面影响准确率等结果,给出固定计算预算下的系统构建指导。

中文摘要 AI 辅助

多页视觉丰富文档理解(MP-VRDU)需要管理稀疏、跨页面分布且常超出模型上下文窗口的证据。现有研究对这类系统的构建方式提出了相互矛盾且大多未经验证的主张。我们将错误答案归因于三种故障模式:表示、选择和推理,并通过固定其余两种模式、干预其中一种模式,在多页文档理解数据集上分离出每种模式的影响。我们发现视觉信息是必要的,但不能替代文本提取;缺失页面会限制准确率,而干扰项的影响很小;即使证据完全提供,推理器也无法整合跨页面的证据。提示能显著改变推理行为,在提升部分结果的同时会损害其他结果。我们将这些发现转化为在固定计算预算下构建此类系统的指导。

英文摘要

Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has produced competing, largely untested claims about how these systems should be built. We attribute incorrect answers to three failure modes, representation, selection, and reasoning, and isolate each over a multi-page document understanding dataset by intervening on one while holding the others fixed. We find that vision is necessary but does not replace text extraction, that missing pages bound accuracy while distractors cost little, and that reasoners fail to integrate evidence across pages even when it is fully supplied. Prompting can shift reasoning behaviour substantially, improving some outcomes at the expense of others. We translate these findings into guidance for building such systems under a fixed compute budget.

发表机构

  • The University of Western Australia(西澳大利亚大学)
  • The University of Melbourne(墨尔本大学)
  • Weill Cornell Medicine(威尔康奈尔医学院)

机构由 AI 辅助整理,请以论文原文为准。

↑