发表机构
University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RidgeRank 通过闭式分数融合和单向量岭回归校正,在视觉文档重排序中实现接近全交叉编码器的精度,并带来高达48倍加速。
AI 中文摘要
多模态语言模型能够准确地对视觉文档检索结果进行重排序,但对每个候选页面以全成本打分使其速度缓慢。一些压缩这些重排序器的方法需要相关性标签来恢复准确性,并且仅依据重排序器分数进行排序。RidgeRank 衡量重排序器分数所缺乏的相关性信号,并通过闭式融合规则从检索器分数中恢复该信号。最大化相关性目标可得到最优融合权重,以及重排序器分数本身无法达到该最优值的精确条件。进一步地,通过一次中心化岭回归,将重排序器应用于未压缩页面的同一模型全深度分数,得到一个单一向量,作用于中间隐藏状态以校正重排序器。在取自 ViDoRe 2 和 ViDoRe 3 的 12 个数据集上,使用两种检索器和两种语言模型骨干进行评估,RidgeRank 将 NDCG@5 提升至与全交叉编码器相差 1.2 个百分点以内,同时实现高达 48 倍的加速,推进了视觉文档重排序的准确性与延迟帕累托前沿。
英文摘要
Multimodal language models rerank visual document retrieval results accurately, but scoring every candidate page at full cost makes them slow. Some methods that compress these rerankers need relevance labels to regain accuracy, and they rank by the reranker score alone. RidgeRank measures how much relevance signal the reranker score lacks and recovers it from the retriever score through a closed-form fusion rule. Maximizing a correlation objective gives the optimal fusion weight, along with the exact condition under which the reranker score by itself cannot reach that optimum. The reranker is further corrected by a single vector applied to an intermediate hidden state, obtained through one centered ridge regression onto the same model's full-depth scores on uncompressed pages. On 12 datasets drawn from ViDoRe 2 and ViDoRe 3, evaluated with two retrievers and two language model backbones, RidgeRank brings NDCG@5 to within 1.2 pp of a full cross encoder with speedups of up to 48 times, advancing the accuracy and latency Pareto frontier for visual document reranking.