ViSAR:用于视觉文档问答的无需训练的自适应k检索方法
ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering
- INSA Lyon(里昂国立应用科学学院)
- CNRS(法国国家科学研究中心)
- EPITA Research Laboratory (LRE)(EPITA研究实验室)
- Lowit(洛维特公司)
- LIRIS UMR 5205(里昂信息、图像与智能系统实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ViSAR是一种无需训练的自适应k检索方法,用于视觉文档问答,可动态确定检索页面数量,降低RAG延迟最高58.7%,同时保持或提升答案准确性。
AI中文摘要:
文档视觉问答(DocVQA)常利用检索增强生成(RAG),其中通常采用晚期交互编码器识别与用户查询相关的文档页面,之后由大型视觉语言模型(LVLM)生成答案。现有方法通常检索固定数量的前k个页面,而不考虑查询的复杂程度,这会增加LVLM的延迟,还可能降低答案准确性。我们提出ViSAR(视觉语义激活检索),一种用于晚期交互视觉文档检索的无需训练的自适应k检索方法。ViSAR直接在嵌入空间中操作,构建查询条件下的页面级相似度矩阵,该矩阵突出显示与查询相关的语义并动态确定要检索的页面数量。在多个编码器和LVLM上,ViSAR检索紧凑的、适配查询的页面集合,与固定前k检索和自适应检索启发式方法相比,它将RAG延迟降低了高达58.7%,同时保持或提高了答案准确性。此外,我们表明相似度矩阵结构与答案准确性相关,这为检索质量感知的文档理解指明了未来方向。
英文摘要:
Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vision-Language Model (LVLM). Existing approaches typically retrieve a fixed top-$k$ number of pages regardless of query complexity, which increases LVLM latency and may degrade answer accuracy. We introduce ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval. ViSAR operates directly in the embedding space to construct a query-conditioned page-level similarity matrix that highlights query-relevant semantics and dynamically determines the number of pages to retrieve. Across multiple encoders and LVLMs, ViSAR retrieves compact, query-adapted page sets that reduce RAG latency by up to 58.7\%, while maintaining or improving answer accuracy compared with fixed top-$k$ and adaptive retrieval heuristics. Furthermore, we show that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.