重新排序前三思:多视角证据与推理融合用于文本重排序
Think Thrice Before Reranking: Multi-perspective Evidence and Reasoning Integration for Text Reranking
浏览论文内容
中文总结 AI 辅助
针对现有LLM重排序依赖单一推理轨迹易出错的问题,提出MERIT-Rank框架,通过多轨迹推理空间与联合重排序器整合多视角证据,并采用渐进式训练优化,在推理密集型及传统基准上取得领先性能。
中文摘要 AI 辅助
基于大语言模型(LLM)的推理式重排序在文本排序中展现出显著的改进效果。然而,当前方法主要依赖单一推理轨迹,导致排序结果易受推理错误影响,并且在建模文档相关性背后的多面信号方面存在固有局限。为解决这一困境,我们提出了MERIT-Rank(多视角证据与推理融合用于文本重排序),一个通过建模互补推理轨迹来提升重排序鲁棒性的框架。MERIT-Rank构建了多轨迹推理空间(MTRS),从多个视角评估查询-文档相关性,并引入一个联合重排序器,将这些推理路径整合为统一的排序决策。我们进一步开发了渐进式排序策略优化(PRPO),一种渐进式训练框架,通过分阶段优化目标稳定推理轨迹并持续提升排序质量。在推理密集型与传统检索基准上的实验表明,MERIT-Rank持续优于竞争基线。值得注意的是,4B模型在BRIGHT基准上显著超越了大多数7B甚至32B的重排序器。
英文摘要
Reasoning-based reranking with Large Language Models (LLMs) has shown promising improvements in text ranking. However, current methods predominantly rely on a single reasoning trajectory, resulting in rankings that are susceptible to reasoning errors and inherently constrained in modeling the multifaceted signals underlying document relevance. To resolve this dilemma, we propose MERIT-Rank(Multi-perspective Evidence and Reasoning Integration for Text Reranking), a framework that models complementary reasoning trajectories to improve reranking robustness. MERIT-Rank formulates a Multi-Trajectory Reasoning Space (MTRS) that evaluates query-document relevance from multiple perspectives and introduces a joint reranker that consolidates these reasoning paths into a unified ranking decision. We further develop Progressive Rank Policy Optimization (PRPO), a progressive training framework that stabilizes reasoning trajectories while continually improving ranking quality through staged optimization objectives. Experiments on both reasoning-intensive and traditional retrieval benchmarks show that MERIT-Rank consistently achieves superior performance over competitive baselines. The 4B model notably outperforms most 7B and even 32B rerankers on BRIGHT.
发表机构
- Honor Device Co., Ltd(荣耀终端有限公司)
机构由 AI 辅助整理,请以论文原文为准。