发表机构
IBM(国际商业机器公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉文档检索的多模态晚期交互检索器存储成本高、压缩级别固定的问题,提出ColSNAP训练方法,可生成嵌套压缩层级,实现灵活的精度-存储权衡,大幅压缩下仍保持检索性能并可迁移至多类骨干网络。
AI 中文摘要
多模态晚期交互检索器通过将每个页面表示为补丁嵌入并在令牌级别进行匹配,在视觉丰富的文档上实现了强大的检索性能,但该方法会产生较高的存储成本。现有压缩方法通常在索引时固定单一压缩级别,灵活性受限。我们提出ColSNAP(空间嵌套平均池化,Spatial Nested Average Pooling),这是一种直接从骨干网络的补丁网格生成嵌套压缩级别的训练方法。通过将补丁嵌入进行空间池化,形成逐渐更粗的层级并同时训练所有层级,单一模型无需改变架构即可支持多压缩级别的检索。关键在于,单次编码即可生成所有层级,使得精度-存储权衡可在索引时根据可用存储预算进行配置,而非在训练时固定。我们证明,使用ColSNAP训练的模型在大幅压缩下仍能保持接近全分辨率的检索性能,且ColSNAP可有效迁移至多款晚期交互骨干网络,其大部分改进通过应用于预训练检索器的轻量适配阶段实现。
英文摘要
Multi-modal late-interaction retrievers achieve strong retrieval on visually rich documents by representing each page as per patch embeddings and matching at the token level. However, this approach incurs high storage costs. Existing compression methods typically fix a single compression level at indexing time, limiting flexibility. We present ColSNAP (Spatial Nested Average Pooling)1, a training method that generates a nested hierarchy of compression levels directly from a backbone's patch grid. By spatially pooling patch embeddings into pro- gressively coarser tiers and training all tiers simultaneously, a single model learns to support retrieval at multiple compression levels without architectural changes. Crucially, a single encoding pass yields every tier, enabling the accuracy-storage trade-off to be configured at indexing time to match avail- able storage budgets, rather than being fixed during training. We demonstrate that models trained using ColSNAP maintain near full-resolution retrieval performance under substantial compression and that ColSNAP transfers effectively across multiple late-interaction backbones, and achieves most of its improvements via a lightweight adaptation stage applied to a pre-trained retriever.