发表机构
Thomson Reuters Labs; University of Toronto; Vector Institute(汤姆森路透实验室; 多伦多大学; 矢量研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现排名质量相当的大语言模型评分器决策存在差异,提出OC-SFT方法减弱顺序依赖性,提升决策稳定性,建议对比研究需报告阈值保留内容和读者回答而非仅排名质量。
AI 中文摘要
重排序器、奖励模型和多文档问答评分器会在一个大语言模型提示中对候选文档或响应进行评分,因此每个评分都依赖于它们的顺序。这类评分器是根据排名质量来选择的,但它们的评分会决定一个决策:评分阈值保留什么、读者回答什么,或者偏好模型选择什么。然而,相同的排名质量并不意味着相同的决策:在段落重排序任务中,5个nDCG@10在0.010以内的训练评分器在重新排序时,保留集的重叠度仅为0.66至0.84;在本研究的对比中,一个已发布的重排序器取得了最高的保留集F1值,但其重叠度仍仅为0.667。我们测试的所有提示时间修改都无法消除这种顺序依赖性:唯一能提升排名质量的修改会使所有三个决策保持不变。顺序一致的监督微调(OC-SFT)通过训练候选评分不依赖于顺序来减弱权重中的顺序依赖性,它保持了排名质量,并在所有三个任务的训练评分器中取得了所有决策稳定性指标的最优表现:它在0.125的排列对上翻转了读者的答案,而其他三个针对顺序的目标的这一比例为0.149至0.164;它在12个基础模型上比顺序平均蒸馏更稳定,且一个OC-SFT排列的保留集重叠度超过了10个现成平均排列的保留集。因此,对比研究应报告阈值保留内容和读者回答内容,而非仅报告排名质量。代码可在该https URL获取。
英文摘要
In passage reranking, response ranking and multi-document question answering, LLMs can score several candidate documents or responses together in one prompt, each still receiving its own score. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, a reader answers, or which chosen/rejected pair enters preference training. Because the candidates share that prompt, reordering them changes their scores. The same query over the same candidates should still yield the same decision. However, equal ranking quality does not imply equal decisions: on passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when reordered. No prompt-time change we test resolves that dependence: the only one that improves ranking quality does not measurably improve decision stability. We introduce order-consistency SFT (OC-SFT), which attenuates it in the weights by penalizing disagreement between a candidate's scores across orderings. It holds ranking quality and leads every decision-stability measure among trained scorers on all three tasks. It is also more stable on 12 base models than order-averaged distillation, which trains on labels averaged across permutations. One OC-SFT permutation retains sets that overlap more than ten averaged off-the-shelf permutations. A comparison of such scorers should therefore report what a threshold retains and a reader answers, not ranking quality alone. Code is available at https://github.com/thomsonreuters/presentation-dependence.
Comments9 pages main text, 45 pages total