发表机构
Ant Group; Tsinghua University(蚂蚁集团; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RandSlot通过随机软标记训练编码器生成紧凑多向量表示,在相同向量预算下提升视觉文档检索质量,且推理时无需采样。
AI 中文摘要
视觉文档检索需要具有表现力的表示,以便将查询与分布在文本、表格和页面布局中的证据进行匹配。多向量表示能够捕获细粒度信息,但存储和比较众多向量会带来可观的检索成本。在本文中,我们介绍了RandSlot,一种利用随机软标记学习紧凑视觉文档表示的简单方法。在训练期间,我们将独立采样的随机单位向量附加到查询和文档输入序列中,并在每次使用时重新采样,而不引入可学习的软标记参数。编码器将这些辅助输入与原始内容进行上下文关联,以生成一小部分检索向量。标准的后期交互目标函数训练编码器在不同的输入条件下提取相关信息。使用不同骨干模型的实验表明,在相同的向量预算下,RandSlot相较于其他读出策略提升了检索质量。进一步的分析表明,当在推理时用零替换随机软标记时,这些增益仍然可以保持,这表明训练期间的随机输入可以改善紧凑的检索表示,即使推理时不再需要采样。
英文摘要
Visual document retrieval requires expressive representations to match queries with evidence distributed across text, tables, and page layouts. Multi-vector representations capture fine-grained information, but storing and comparing many vectors introduces substantial retrieval costs. In this paper, we introduce RandSlot, a simple approach to learn compact visual document representations with random soft tokens. During training, we append independently-sampled random unit vectors to query and document input sequences and resample them at every use, without introducing learnable soft-token parameters. The encoder contextualizes these auxiliary inputs with the original content to produce a small set of retrieval vectors. A standard late-interaction objective trains the encoder to extract relevant information under varying input conditions. Experiments with different backbone models show that RandSlot improves retrieval quality over alternative readout strategies under the same vector budget. Further analysis shows that these gains can persist when random soft tokens are replaced with zeros at inference, demonstrating that random inputs during training can improve compact retrieval representations even when inference no longer requires sampling.