arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30904cs.IR

QReason:面向查询的解耦思维链用于高效段落重排序

QReason: Query-Focused Decoupled Chain-of-Thought for Efficient Passage Reranking

发表机构北京师范大学人工智能学院 · 北京市教育人工智能重点实验室 · 教育部智能技术与教育应用工程研究中心
另 2 家 · 查看机构详情
  • School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
  • Beijing Key Laboratory of Artificial Intelligence for Education(北京市教育人工智能重点实验室)
  • Engineering Research Center of Intelligent Technology and Educational ApplicationMinistry of Education(教育部智能技术与教育应用工程研究中心)
  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

Yang Zhang, Wenhan Liu, Qiannan Zhu, Mingming Li, Yuanfei Huang

首次发表
浏览论文内容

中文总结 AI 辅助

QReason提出解耦框架,将查询聚焦推理与窗口评估分离,通过两阶段训练生成可复用推理链,在BRIGHT基准上减少冗余并提升重排序效率与性能。

中文摘要 AI 辅助

段落重排序通过优化候选段落的排序以更好地反映相关性,在信息检索中发挥着关键作用。现有的基于思维链(CoT)推理的列表式大语言模型(LLM)重排序器能够有效处理复杂查询,但由于滑动窗口策略会重复生成高度相似的思维链,导致大量冗余和高延迟。为解决这一问题,我们提出了QReason,一个将面向查询的推理与特定窗口的段落相关性评估相分离的解耦框架。具体而言,QReason引入了一个专门的改写器,仅生成一次面向排序的推理查询,捕获查询的核心意图并避免冗余推理,然后在所有窗口中使用一个不进行推理的重排序器复用该查询。该改写器通过两阶段流程进行训练:首先使用带相关段落指导的监督微调,通过语义证据生成深度扎根的、面向查询的思维链;然后应用强化学习,使思维链生成与推理时设置和重排序目标对齐,优化列表式指标和段落级区分度,以生成可复用的推理链用于重排序。在BRIGHT基准上的实验表明,QReason显著减少了冗余推理,取得了与强推理型重排序器相当或更优的排序性能,并优于现有的查询改写模型。

英文摘要

Passage reranking plays a crucial role in information retrieval by refining the ordering of candidate passages to better reflect relevance. Existing listwise LLM rerankers with Chain-of-Thought (CoT) reasoning can handle complex queries effectively, but they suffer from substantial redundancy and high latency due to sliding-window strategies, which repeatedly generate highly similar CoTs. To address this, we propose QReason, a decoupled framework that separates query-focused reasoning from window-specific passage relevance assessment. Specifically, QReason introduces a dedicated rewriter that generates a ranking-oriented reasoning query once, capturing the query's core intent while avoiding redundant reasoning, and then reuses it across all windows with a non-reasoning reranker. The rewriter is trained via a two-stage process that first uses supervised fine-tuning with relevant-passage guidance through semantic evidence to produce deeply grounded, query-focused CoTs. It then applies reinforcement learning to align CoT generation with both the inference-time setting and the reranking objective, optimizing listwise metrics and passage-level discrimination to produce reusable reasoning chains for reranking. Experiments on the BRIGHT benchmark demonstrate that QReason significantly reduces redundant reasoning, achieves ranking performance comparable to or better than strong reasoning-based rerankers, and outperforms existing query rewriting models.

补充信息

↑