arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TestHallVQA:探索LVLM在科学考试冗余上下文中的文档级推理

TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams

Yongqi Yu, Yu Zhang

arXiv 2609.13158首次发表:更新:

发表机构

Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出TestHallVQA基准,结合文档级规模与考试难度,并引入F1-R²指标,以评估LVLM在冗余上下文中的推理与证据检索能力,揭示其潜在缺陷。

AI 中文摘要

大型视觉-语言模型(LVLMs)日益被期望对平面媒体执行视觉问答(VQA)。然而,现有的平面VQA基准通常强调孤立的挑战:有些强调长文档理解但推理深度有限,而另一些则需要复杂的视觉推理但仅限于单页、无噪声的设置。此外,通过理论分析,我们识别了无关视觉标记的影响,这导致可测量的性能下降,但在系统性量化方面却很少受到关注。为解决这些局限性,我们引入了TestHallVQA,一个多图像VQA基准,同时体现了文档级规模与人类考试的难度,并提供全面的任务覆盖。利用TestHallVQA可控注入多级上下文冗余的能力,我们进一步提出了一种新颖的指标F1-R²,该指标联合量化了LVLMs的计算推理能力及其针对文档级冗余的证据检索鲁棒性。对主流LVLMs的大量实验和分析揭示了它们在多个维度上的潜在缺陷,为未来研究提供了具体见解和方向。相关的数据集、代码和完整的理论推导可在以下https URL获取。

英文摘要

Large Vision--Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar VQA benchmarks typically emphasize isolated challenges: some emphasize long-document understanding with limited reasoning depth, while others require complex visual reasoning but remain restricted to single-page, noise-free settings. Moreover, through theoretical analysis, we identify the impact of irrelevant visual tokens, which leads to measurable performance degradation but has received little attention with respect to systematic quantification. To address these limitations, we introduce TestHallVQA, a multi-image VQA benchmark that simultaneously embodies document-level scale and the difficulty of human examinations, while providing comprehensive task coverage. Leveraging TestHallVQA's ability to controllably inject multi-level contextual redundancy, we further propose a novel metric, F1-R\textsuperscript{2}, which jointly quantifies LVLMs' computational reasoning capability and their evidence retrieval robustness against document-level redundancy. Extensive experiments and analyses on mainstream LVLMs reveal their latent deficiencies across multiple dimensions, offering concrete insights and directions for future research. The associated datasets, code, and complete theoretical derivations are available at https://github.com/yqyu2317/TestHallVQA-benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑