DistilVDR:通过双学生蒸馏实现的紧凑端到端视觉文档检索器
DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
浏览论文内容
中文总结 AI 辅助
本研究提出DistilVDR,通过双学生蒸馏从80亿参数视觉-语言教师模型得到5.24亿参数的紧凑端到端VDR系统,在多个基准上性能接近教师模型且索引速度更快、体积更小。
中文摘要 AI 辅助
视觉文档检索(VDR)目前由数十亿参数的模型主导,这类模型在全语料库规模下索引速度慢,服务成本高。现有的压缩方案要么从头训练更小的多向量编码器,要么仅对查询侧进行蒸馏,均无法得到紧凑的单向量端到端检索器。我们提出DistilVDR,这是一个5.24亿参数的端到端VDR系统,通过在逐点余弦对齐损失下从单个80亿参数的视觉-语言教师模型进行双侧蒸馏得到。所有监督信号均来自冻结教师模型的嵌入空间,该空间本身已使用相关性监督进行训练,因此学生模型的目标无需相关性标签、负采样或对比项。我们通过仅编码器的学生模型来适配VDR的文本查询与图像文档输入的不对称性,该模型将视觉能力集中在文档侧,查询侧保持7000万参数。我们发布了两个变体,它们共享相同的编码器和训练过程,仅在文档编码器的视觉块预算上不同:DistilVDR-HiRes在ViDoRe v1+v2+v3数据集上达到61.74的平均NDCG@5(相当于80亿参数教师模型性能的86.9%),并在对高分辨率敏感的v3基准上优于所有复现的10亿参数以下的基线模型;而DistilVDR-Fast在视觉令牌预算缩小3倍的情况下达到59.98的性能。两个变体存储100万份文档的索引,比最强的10亿参数以下多向量基线小15.6倍,且语料库索引速度快一个数量级。代码可在此https URL获取。
英文摘要
Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Prior compression routes either train a smaller multi-vector encoder from scratch or distil only the query side; neither yields a compact single-vector retriever end-to-end. We present DistilVDR, a 524M end-to-end VDR system distilled bilaterally from a single 8B vision-language teacher under a pointwise cosine alignment loss. All supervision comes from the frozen teacher's embedding space, which was itself trained with relevance supervision, so the student objective needs no relevance labels, negative sampling, or contrastive term. We match VDR's text-query and image-document input asymmetry with an asymmetric encoder-only student that concentrates visual capacity on the document side and keeps the query side at 70M parameters. We release two variants that share the same encoders and training and differ only in the document encoder's visual-tile budget: DistilVDR-HiRes attains 61.74 average NDCG@5 on ViDoRe v1+v2+v3 (86.9% of the 8B teacher) and leads every reproduced sub-1B baseline on the high-resolution-sensitive v3 benchmark, while DistilVDR-Fast attains 59.98 at a 3 times smaller visual-token budget. Both variants store one million documents in a 15.6 times smaller index than the strongest sub-1B multi-vector baseline and index the corpus an order of magnitude faster. The code is available at https://github.com/Ryenhails/NanoVDR.
发表机构
- Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。