arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34899cs.IRcs.CLcs.CV

ColNanoVDR:基于最优传输的无文档查询蒸馏用于多向量视觉文档检索

ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport

Zhuchenyang Liu, Ziyi Wang, Yao Zhang, Yu Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

ColNanoVDR提出首个无文档蒸馏框架,通过带学习权重的最优传输对齐多向量视觉文档检索中教师与学生模型的查询令牌,无需页面编码,在ViDoRe基准上保留95%性能且查询提速26倍。

中文摘要 AI 辅助

基于视觉语言模型的多向量检索器引领了视觉文档检索(VDR)领域,但它们在每次搜索时都需要运行一个拥有数十亿参数的查询编码器。将该编码器蒸馏为一个小型学生模型,使其能够查询教师模型已有的索引,将消除这一瓶颈。然而,标准方法需要匹配教师模型的MaxSim分数,因此需要对每个训练页面进行编码和缓存,这可能导致页面令牌达到TB级别。NanoVDR完全避免了页面处理,仅通过教师模型的查询嵌入进行训练,但该方法仅适用于单向量检索器。我们提出了ColNanoVDR,据我们所知,这是首个将这种无文档蒸馏引入多向量VDR的框架。其目标函数OTW(带学习权重的最优传输)通过熵正则化最优传输将学生模型的查询令牌与教师模型对齐,并为每个学生令牌学习一个权重,且不需要两种令牌化之间的对应关系。我们证明了由此产生的对齐成本在每一页上界定了MaxSim分数差异。从五个最先进的教师模型蒸馏得到的1.49亿参数纯文本学生模型,在ViDoRe v1-v3上保留了教师模型约95%的NDCG@5,同时查询编码速度提升高达26倍。在相同训练条件下,OTW在无需编码任何页面且读取的缓存教师数据减少12.6倍的情况下,达到了与分数蒸馏相当的性能。

英文摘要

Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard recipe, however, matches the teacher's MaxSim scores and so requires encoding and caching every training page, which can reach terabytes of page tokens. NanoVDR avoids pages entirely by training on the teacher's query embeddings alone, but only for single-vector retrievers. We present ColNanoVDR, to our knowledge the first framework to bring this document-free distillation to multi-vector VDR. Its objective, OTW (Optimal Transport with Learned Weights), aligns the student's query tokens with the teacher's by entropic optimal transport, with a learned weight for each student token, and needs no correspondence between the two tokenizations. We prove that the resulting alignment cost bounds the MaxSim score difference on every page. Distilled from five state-of-the-art teachers, the 149M text-only students retain about 95% of their teachers' NDCG@5 on ViDoRe v1-v3 while encoding queries up to 26x faster. Under identical training, OTW matches score distillation while encoding no page and reading 12.6x less cached teacher data.

发表机构

  • Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑