发表机构
State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院通用人工智能省部共建国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出VisDocAgentBench基准,对比静态与智能体检索,发现视觉检索优于OCR文本检索,智能体可改善双桥项检索性能,为检索智能体发展提供方向。
AI 中文摘要
视觉丰富文档通过语言、布局、结构化视觉元素和语料库上下文编码相关性,但检索通常通过单次查询-页面匹配进行评估。智能体搜索基准通常对下游问答或报告生成进行评分,而迭代证据获取下的文档排名却未被充分探索。我们引入VisDocAgentBench,这是一个封闭语料库基准,在共享排名输出协议下比较静态检索和智能体检索。它包含来自100份文档的2375个页面,以及120个唯一目标查询,这些查询在直接、单桥和双桥证据结构间平衡分配。保持关系的构建方式产生了语义、关系和视觉查询,随后进行全文档审查和难负样本验证。一个强大的后期交互视觉检索器在直接项上达到97.50%的Recall@1,但在双桥项上仅为2.50%,这暴露了当相关性依赖于语料库上下文时,查询-目标匹配的局限性。智能体挽回了大部分损失,但规划器选择和检索表示仍然是决定性因素。所有规划器在视觉检索下表现更好,其最佳R@1达到67.50%,而OCR文本检索为37.50%。消融实验确定迭代搜索和页面检查是关键能力,提供完整支持上下文可改善两种路径的排名。轨迹分析将剩余损失定位在目标发现、候选检查和证据角色整合上。这些发现推动了结合模态保留发现与证据导向验证的检索智能体的发展。
英文摘要
Visually rich documents encode relevance through language, layout, structured visual elements, and corpus context, yet retrieval is typically evaluated by one-shot query--page matching. Agentic-search benchmarks usually score downstream question answering or report generation, leaving document ranking under iterative evidence acquisition underexplored. We introduce VisDocAgentBench, a closed-corpus benchmark comparing static and agentic retrieval under a shared ranked-output contract. It contains 2,375 pages from 100 documents and 120 unique-target queries balanced across direct, one-bridge, and two-bridge evidence structures. Relation-preserving construction yields semantic, relational, and visual queries, followed by full-document review and hard-negative validation. A strong late-interaction visual retriever reaches 97.50% Recall@1 on direct items but 2.50% on two-bridge items, exposing the limits of query--target matching when relevance depends on corpus context. Agents recover much of this loss, but planner choice and retrieval representation remain decisive. Every planner performs better with visual retrieval, whose best R@1 reaches 67.50% versus 37.50% for OCR-text. Ablations identify iterative search and page inspection as consequential capabilities, and providing the complete support context improves ranking on both routes. Trace analysis localizes the remaining losses to target discovery, candidate examination, and evidence-role integration. These findings motivate retrieval agents that combine modality-preserving discovery with evidence-directed verification.