arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DocHop:面向信息密集型文档的域外多跳推理基准测试

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park, Xinyi Gu, Zexue He, Soochahn Lee, Rogerio Feris, Yong Jae Lee

arXiv 2609.02059首次发表:更新:

发表机构

University of Wisconsin-Madison; Massachusetts Institute of Technology; Stanford University; Kookmin University; MIT-IBM Watson AI Lab, IBM Research(威斯康星大学麦迪逊分校; 麻省理工学院; 斯坦福大学; 国民大学; 麻省理工学院-IBM沃森人工智能实验室,IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出DocHop基准,针对信息密集型文档的域外多跳推理,构建含2074个样本的测试集,发现MLLMs性能远低于人类,为相关研究提供可控测试平台。

AI 中文摘要

多模态大语言模型(Multimodal Large Language Models, MLLMs)在图表和文档问答等结构化视觉理解任务上已取得出色性能。然而,现有基准通常孤立评估这些领域,忽略了一项关键能力:模型能否利用文本上下文确定如何选择、解读和整合图表证据。我们推出DocHop——一个针对文档风格图像中图表-上下文集成推理的基准。在DocHop中,文档叙述指定多步组合约束,图表提供对应数据值;问题基于叙述中定义的语义参考标签,要求模型从上下文中解析目标实体,再整合多张图表的证据。为实现系统评估,我们采用随机逻辑优先生成流程构建DocHop,可控制推理深度和视觉密度,涵盖6个任务类别的2074个样本。对多种专有及开源MLLMs的实验显示,其与人类性能存在显著差距:标注者准确率超90%,最优模型仅达62.83%;推理增强模型性能始终有所提升,但随推理复杂度增加而下降。总体而言,DocHop为挑战性多跳文档推理提供了可控测试平台。

英文摘要

Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.

CommentsAccepted by ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑