arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35069cs.IR

RenderRank:利用压缩视觉令牌学习文本重排序

RenderRank: Learning to Rerank Text with Compressed Visual Tokens

Seongtae Hong, Youngjoon Jang, Jungseob Lee, Hyeonseok Moon, Heuiseok Lim

首次发表 更新
浏览论文内容

中文总结 AI 辅助

该研究提出重排序模型RenderRank,通过将文档渲染为图像并编码为压缩视觉令牌实现相关性打分,在减少输入令牌的同时,在BEIR及长文档数据集上取得优于多数文本基线的性能与更高吞吐量,为重排序提供了新的表示方案。

中文摘要 AI 辅助

将文档文本渲染为图像可使视觉-语言模型将文档编码为视觉令牌,与文本输入相比能够缩短输入序列长度。这种输入长度的缩减在重排序任务中尤为实用——重排序中每个查询都需要为多个候选文档打分,令牌节省效应可作用于每一次候选评估过程。我们提出RenderRank,这是一种重排序器,它从压缩的视觉文档表示中学习与查询相关的相关性打分,而非采用传统基于文本的重排序器所使用的文本令牌序列。其训练流程首先将视觉输入的相关性得分与基于文本的教师模型的得分进行对齐,随后针对同一查询优化正例文档和负例文档的相对得分。在BEIR的11个数据集上,RenderRank的输入令牌数量减少了16.5%-35.5%,同时平均NDCG@10达到55.96,性能优于所有参数量低于4B的受评估文本基线模型,也超过了部分更大规模的模型。在4个长文档数据集上,它的平均NDCG@10达到88.27,而平均输入令牌数约为受评估文本重排序器的一半。在该场景下,RenderRank的平均吞吐量是受评估基线中最高值的1.70倍。这些结果表明,压缩视觉表示能够支撑准确的文档相关性打分,为重排序任务提供了文本令牌表示之外的替代方案。

英文摘要

Rendering document text as images allows vision-language models to encode documents as visual tokens, which can reduce input sequence length compared with text input. This reduction in input length is particularly useful for reranking, where each query involves scoring multiple candidate documents and token savings apply to each candidate evaluation. We introduce RenderRank, a reranker that learns query-dependent relevance scoring from compressed visual document representations instead of the text token sequences used by conventional text-based rerankers. Training first aligns relevance scores from visual inputs with those of a text-based teacher, then refines the relative scores of positive and negative documents for the same query. Across 11 datasets from BEIR, RenderRank uses 16.5-35.5% fewer input tokens while achieving an average NDCG@10 of 55.96, outperforming all evaluated text-based baselines below 4B parameters and some larger models. Across four long-document datasets, it achieves an average NDCG@10 of 88.27 with approximately half the average input token count of the evaluated text-based rerankers. In this setting, RenderRank delivers 1.70x the highest average throughput of the evaluated baselines. These results demonstrate that compressed visual representations can support accurate document relevance scoring, providing an alternative to text token representations for reranking.

发表机构

  • Korea University(高丽大学)
  • Sookmyung Women’s University(淑明女子大学)

机构由 AI 辅助整理,请以论文原文为准。

↑