arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

查询感知的令牌预算分配用于高效晚期交互视觉文档检索

Query-Aware Token Budgeting for Efficient Late-Interaction Visual Document Retrieval

PS Rishi, Rajeev Ranjan Dwivedi, Vinod K Kurmi

arXiv 2609.07262首次发表:更新:

发表机构

Indian Institute of Science Education and Research Bhopal (IISER Bhopal)(印度科学教育与研究学院博帕尔分院(IISER 博帕尔))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对晚期交互视觉文档检索中静态池化损失证据的问题,提出查询感知的令牌预算分配,通过有预算的MaxSim覆盖选择令牌,在ViDoRe任务上恢复98.39%的检索质量,优于静态池化。

AI 中文摘要

晚期交互视觉文档检索器通过每页存储大量令牌嵌入来保留细粒度的页面证据,但由此产生的存储和查询时交互成本使得大规模部署变得昂贵。在索引前对文档令牌进行池化提供了一种自然的补救措施,然而静态池化必须在查询已知之前决定保留哪些视觉证据。我们研究了一种替代方案:一个高度压缩的热路径索引生成候选结果,之后查询感知的令牌预算分配作用于入围页面的原始令牌集。我们将这一第二阶段选择形式化为一个有预算的MaxSim覆盖问题,证明其截断版本是单调子模的,并比较了仅覆盖、聚类引导、逐令牌和边际增益策略。在十个ViDoRe任务上使用ColModernVBERT,直接静态池化将宏平均归一化折现累积增益(第5位)从无压缩时的0.6309降至三十二倍池化因子下的0.4738。在相同的候选生成机制和相当于八倍池化因子的重排序预算下,令牌top-k恢复了全令牌分数的93.93%,而贪心边际增益选择恢复了98.39%。留出法和逐数据集留一法评估表明,贪心方法在每个数据集上相对于令牌top-k都有正向改进。延迟分析揭示了两个有用的操作点:令牌top-k适用于交互式检索,而朴素贪心实现则作为质量上界。这些结果共同表明,晚期交互视觉检索受益于查询感知的分配,而非仅依赖查询无关的池化。

英文摘要

Late-interaction visual document retrievers preserve fine-grained page evidence by storing many token embeddings per page, but the resulting storage and query-time interaction costs make large-scale deployment expensive. Pooling document tokens before indexing offers a natural remedy, yet static pooling must decide which visual evidence to preserve before the query is known. We study an alternative: a heavily compressed hot-path index generates candidates, after which query-aware token budgeting operates on the original token sets of the shortlisted pages. We formulate this stage-two selection as a budgeted MaxSim coverage problem, show that a clipped version is monotone submodular, and compare coverage-only, cluster-guided, token-wise, and marginal-gain policies. On ten ViDoRe tasks with ColModernVBERT, direct static pooling reduces macro normalized discounted cumulative gain at rank five from 0.6309 without compression to 0.4738 at a thirty-two-fold pool factor. Under the same candidate-generation regime and a pool-factor-eight-equivalent reranking budget, token top-k recovers 93.93 percent of the full-token score, while greedy marginal-gain selection recovers 98.39 percent. Held-out and leave-one-dataset-out evaluations yield positive greedy improvements over token top-k on every dataset. The latency analysis reveals two useful operating points: token top-k for interactive retrieval and the naive greedy implementation as a quality upper envelope. Together, these results show that late-interaction visual retrieval benefits from query-aware allocation rather than query-agnostic pooling alone.

CommentsAccepted at IEEE International Conference on Data Mining (ICDM)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑