arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VaRS-Doc:面向视觉文档检索的、通过潜在自探测实现的、具备解释感知能力的变体表示方法

VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval

Haocheng Wang, Tongkun Guan, Wei Shen, Xiaokang Yang

arXiv 2608.01211首次发表:更新:

AI 中文总结

本文针对视觉文档检索中固定文档表示难以适配不同查询意图的问题,提出VaRS-Doc框架,通过两阶段训练实现变体表示,在基准测试中达到最先进性能,解决了查询与文档编码的不匹配问题。

AI 中文摘要

视觉文档检索在企业搜索、科学文献发现、检索增强生成等应用中愈发重要,这类应用依赖于在大型视觉丰富文档集合中高效识别与查询相关的页面。现有方法通常采用后期交互架构,该架构会对文档进行离线编码与索引,以实现可扩展、低延迟的在线检索。尽管该范式效率较高,但它需要在查询已知前将每个文档编码为固定表示;然而,视觉文档中的相同内容可能会因查询意图不同而产生不同的解释,固定表示难以捕捉这一点,若将文档编码推迟至查询到达时进行,则会导致在线检索延迟过高。为解决这一差距,本文提出VaRS-Doc,这是一种视觉文档检索框架,该框架通过让模型在文档编码过程中主动探索变体潜在解释,使文档表示多样化,同时保留高效的后期交互检索,其中每个查询可自适应选择最适配的表示。本文进一步引入两阶段训练策略,鼓励模型捕捉互补的语义解释,并防止其退化为训练单一主导表示。在视觉文档检索基准上的实验表明,VaRS-Doc实现了最先进的检索性能,为解决与查询无关的文档编码和与查询特定的检索需求之间的不匹配问题提供了实用解决方案。代码可在this https URL获取。

英文摘要

Visual document retrieval has recently become increasingly important in applications such as enterprise search, scientific literature discovery, and retrieval-augmented generation. These applications depend on efficiently identifying query-relevant pages across large collections of visually rich documents. Existing methods commonly adopt late-interaction architectures that encode and index documents offline to enable scalable and low-latency online retrieval. Despite its efficiency, this paradigm requires each document to be encoded into a fixed representation before the query is known. However, the same content in a visual document may induce different interpretations depending on the query intent, which a fixed representation struggles to capture. Yet postponing document encoding until the query arrives would incur prohibitive online retrieval latency. To address this gap, we propose VaRS-Doc, a visual document retrieval framework that diversifies document representations by enabling the model to actively explore variant latent interpretations during document encoding, while preserving efficient late-interaction retrieval in which each query adaptively selects the best-fit representation. We further introduce a two-stage training strategy that encourages the model to capture complementary semantic interpretations and prevents it from falling back to train a single dominant representation. Experiments on visual document retrieval benchmarks show that VaRS-Doc achieves state-of-the-art retrieval performance, offering a practical solution to the mismatch between query-agnostic document encoding and query-specific retrieval needs. Code is available at https://github.com/bokufa/VaRS-Doc.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑