arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SynthDocBench:用于长上下文视觉文档理解的可控基准测试

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

Abhigya Verma, Khyati Mahajan, Amit Kumar Saha, Shruthan Radhakrishna, Sagar Davasam, Vikas Yadav, Sai Rajeswar Mudumba

arXiv 2607.10400首次发表:更新:

AI 中文总结

研究针对视觉语言模型在现实文档理解中归因失败原因困难的问题,引入SynthDocBench基准测试,通过组合设计控制多因素,端到端生成多样长文档,评估模型发现三种现有基准未暴露的失败模式,揭示模型或过度拟合基准工件。

AI 中文摘要

视觉语言模型在视觉文档理解基准测试中表现出色,但现实世界文档包含多种因素,难以确定模型失败原因。我们引入SynthDocBench,这是一个用于长上下文视觉文档理解的完全合成基准测试,系统控制文档长度、布局结构、模态组成和问题类型等因素。通过组合设计构建基准测试,各因素独立变化。文档通过LLM管道端到端生成,跨越六种布局原型,40%随机覆盖防止模型利用虚假关联。该基准测试文档更长、结构更多样。评估七个前沿视觉语言模型,发现现有基准测试未暴露的三种失败模式,表明当前模型可能过度拟合基准测试工件而非实现强大的长上下文视觉文档理解。

英文摘要

Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constructed using a combinatorial design, each factor is varied independently across generated documents, enabling controlled analysis of model behavior. Documents are generated end to end using an LLM pipeline across six layout archetypes, with a 40 percent random override to prevent models from exploiting spurious correlations. Additionally, SynthDocBench spans long-context documents with substantially greater length and structural diversity than existing benchmarks. Evaluating seven frontier VLMs, we uncover three failure modes that existing benchmarks cannot surface: sharp degradation with document length, a systematic positional sensitivity in which the middle third of a document is hardest for five of six models and five of six models show a negative Early-to-Late trend (steepest decline: 8.3 percentage points), and breakdown of chart comprehension in long-document settings. These results suggest that current models may be overfitting to benchmark artifacts rather than achieving robust long-context visual document understanding.

Comments29 Pages, 27 Tables, 13 Figures, Accepted at COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑