arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaTutoRank:通过自适应辅导优化学习对文档集合进行重排序,用于RAG与深度研究

AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research

Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi

arXiv 2609.32472首次发表:更新:

AI 中文总结

针对RAG和深度研究中文档重排序监督稀疏的问题,提出AdaTutoRank,利用自适应辅导优化在三级评分层次下提供银标签、奖励和提示,实现密集监督,在十个基准上取得最优性能并减少检索调用。

AI 中文摘要

文档重排序器决定了在RAG和深度研究中哪些证据能够到达下游模型,然而主流重排序器依据相关性匹配进行选择,而单独相关的文档很少能构成复杂信息需求所要求的完整、互补且无冗余的集合。先前的工作以集合的总体评分分数作为奖励,将目标从对文档排序转变为组合集合。然而,该分数是集合中每篇文档共享的一个标量,因此监督是稀疏的:当集合得分良好时,冗余文档会随其余文档一起获得奖励;当集合得分不佳时,关键文档会随其余文档一起受到惩罚;信用分配使得贡献者与搭便车者难以区分。在线策略蒸馏可以使这种监督变得密集,但现有方法给每次轨迹提供相同的固定指导,对强轨迹而言过于规定性,对弱轨迹而言又过于抽象。因此,我们提出AdaTutoRank,一种使用自适应辅导优化(ATO)在九个评分维度的三级层次结构下训练的集合级重排序器,它为冷启动提供银标签,为强化学习提供奖励,并为蒸馏提供提示。ATO从策略自身的冻结快照中提取三种特异性递增的提示形式:仅评分规则、自选择器在评分规则下选择的兄弟集合,以及自反思器对比轨迹与该兄弟集合的反思;每个轨迹接收与其质量匹配的提示形式。在提示条件化的冻结教师模型和无提示快照下重新评分该轨迹,将提示的效果蒸馏为令牌级优势,该优势补充了组相对结果优势。在涵盖RAG、深度研究和集合级评估的十个基准上,AdaTutoRank在发出更少检索调用的情况下取得了最佳的整体性能。

英文摘要

Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work rewards a set by its aggregate rubric score, shifting the objective from ranking documents to composing sets. Yet that score is one scalar shared by every document in the set, so the supervision is sparse: a redundant document is rewarded with the rest whenever the set scores well, and a decisive one penalized with the rest whenever it does not; credit assignment leaves contributors indistinguishable from free riders. On-policy distillation could densify this supervision, but existing methods give every rollout the same fixed guidance, too prescriptive for strong rollouts and too abstract for weak ones. We therefore propose AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation. ATO draws three hint forms of increasing specificity from the policy's own frozen snapshot: the rubrics alone, a self-selector's sibling-set chosen under rubrics, and a self-reflector's reflection contrasting the rollout with that sibling-set; each rollout receives the form matched to its quality. Re-scoring that rollout under the hint-conditioned frozen teacher and the hint-free snapshot distills the hint's effect into a token-level advantage that complements the group-relative outcome advantage. Across ten benchmarks spanning RAG, deep research, and setwise evaluation, AdaTutoRank attains the best overall performance while issuing fewer retrieval calls.

CommentsProject Page: https://adatutorank.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑