arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从VLM教师中迁移什么?比较视觉文档检索的监督信号

What Transfers from a VLM Teacher? Comparing Supervision Signals for Visual Document Retrieval

Saba Sturua, Han Xiao

arXiv 2610.09177首次发表:更新:

发表机构

Jina AI by Elastic(Elastic旗下Jina AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究比较VLM教师监督信号,发现用教师评判难负样本优于丰富正样本,提升ViDoRe检索性能,并验证其可靠性。

AI 中文摘要

视觉文档检索器通过对比训练:每个查询匹配到一个标记为相关(即正样本)的页面,并推离负样本(即假定不相关的页面)。近期方法通过丰富该正样本,将视觉语言模型(VLM)教师的注意力或描述迁移到检索器中。我们探究教师是否更适合用于另一侧,即评判检索器挖掘为负样本的候选页面,这些页面标签未提供任何信息。在固定学生、数据、优化器和评估的情况下,教师评判的难负样本和分数蒸馏将ViDoRe v2的nDCG@5从55.2提升至62.6和63.0;描述对齐(在此处调整)提升2.6分,注意力锚定无显著提升。与无教师规则(在相同训练计算下从同一挖掘池中选择四个候选,其中最佳为当前系统使用的正样本感知阈值)相比,教师的评判在v2上增加4.1分,在v3上增加1.7分。这与标签的不完整性一致。标注者判断查询的四个排名最高的挖掘候选中有约两个相关,但均未标记,因此训练将检索器推离被当作负样本的相关页面。到达学生的信息是粗糙的:在贪心解码的0-100评分提示下,82%的教师评分落在刻度的一端或另一端,而相关/不相关划分保留了大部分蒸馏收益。十名标注者的审计将教师置于人类标注者之间的变异范围内,并发现其在查询有单一确定答案时可靠。我们在此https URL发布代码、教师的330万条评判和页面描述、挖掘池、人类审计及训练好的适配器。

英文摘要

Visual document retrievers are trained contrastively: each query is matched to one page labelled relevant - the positive - and pushed away from negatives, pages presumed irrelevant. Recent methods distil a vision-language model (VLM) teacher into the retriever by enriching that positive, transferring the teacher's attention over it or a description of it. We ask whether the teacher is better spent on the other side, judging the candidates the retriever mines as negatives, which the label says nothing about. With student, data, optimizer and evaluation fixed, teacher-judged hard negatives and score distillation raise ViDoRe v2 nDCG@5 from 55.2 to 62.6 and 63.0; description alignment, as adapted here, gains 2.6 points and attention grounding nothing measurable. Against teacher-free rules that select four candidates from the same mined pool at identical training compute, the best of which is the positive-aware threshold current systems use, the teacher's judgement adds 4.1 points on v2 and 1.7 on v3. This is consistent with how incomplete the labels are. Annotators judge about two of a query's four top-ranked mined candidates relevant, none of them labelled, so training pushes the retriever away from relevant pages treated as negatives. What reaches the student is coarse: under a greedily decoded 0-100 rating prompt, 82% of the teacher's ratings come back at one end of the scale or the other, and a relevant/irrelevant partition keeps most of the distillation gain. A ten-annotator audit places the teacher within the range of variation among human annotators, and finds it reliable where a query has a single determinate answer. We release the code, the teacher's 3.3M judgements and page descriptions, the mined pools, the human audit and the trained adapters at https://github.com/elastic/vdr-teacher-signals.

Comments27 pages, 1 figure, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑