arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35739cs.IR

Rubric-Calibrated Preferences:基于项目反应理论的大语言模型判断跨查询校准

Rubric-Calibrated Preferences: Cross-Query Calibration of LLM Judgments via Item Response Theory

Fabian David Schmidt, Donato Crisostomi, Carlos Lassance, Nils Reimers

首次发表 更新
浏览论文内容

中文总结 AI 辅助

针对重排器评估指标nDCG依赖离散人工标签、难区分相近质量重排器的问题,提出结合相对与绝对判断、用IRT跨查询校准的RCP方法,其RCP-nDCG指标与人类判断相关性更高,能区分更多重排器对。

中文摘要 AI 辅助

重排器决定用户和大语言模型(LLM)能看到哪些文档,但其标准指标nDCG依赖的人工相关性标签成本高、稀疏、有噪声且为离散分级。随着重排器的质量逐渐接近,基于这些标签计算的nDCG越来越难以区分它们的优劣。LLM评判器可以提供密集标签。单个查询内的相对判断甚至能区分表现接近的候选结果,但不同查询的分数没有统一的尺度。绝对评分有统一尺度,但粒度太粗,无法区分相关性相近的文档。我们提出Rubric-Calibrated Preferences(RCP,评分量规校准偏好)方法,结合了这两种判断方式。该方法通过列表式Bradley-Terry锦标赛对每个查询的文档排序,再由一套严格程度递增的是/否标准构成的评分量规提供绝对基准。项目反应理论(IRT)是一种根据受试者对公共问题的回答来评分的方法,它利用共享的标准将所有查询的锦标赛分数映射到统一尺度上。RCP的检索指标RCP-nDCG用校准后得到的相关性概率替换了nDCG的离散标签。对照46名外部标注者的盲评结果,校准将查询平均得分与人工平均评分之间的相关性从0.538提升至0.795。该概率将有用文档排在无用文档之前的概率为0.910(AUC值,随机水平为0.5),而基准标签的这一数值为0.651。当标注者的评分更偏好两个重排器中的一个,且恰好有一个指标与该偏好一致时,在185组比较中有72.4%的情况里这个指标是RCP-nDCG(随机概率约为53%)。在TREC-DL数据集上,对于所有NIST评估者的评分能显著区分的重排器对,RCP-nDCG的判断都与评估者评分一致。RCP-nDCG还解决了nDCG的许多平局问题,在NanoBEIR数据集上能区分的重排器对数量是nDCG的1.9倍。因此,评分量规校准将相对的LLM判断转化为可跨查询比较、且与人类判断一致的密集相关性标签。

英文摘要

Rerankers decide which documents users and LLMs see, yet their standard metric, nDCG, relies on human relevance labels that are costly, sparse, noisy, and discretely graded. As rerankers approach each other in quality, nDCG on these labels therefore increasingly fails to separate them. LLM judges could supply dense labels. Relative judgments within one query tell even close candidates apart, yet their scores share no scale across queries. Absolute grades share one scale but are too coarse to distinguish documents of similar relevance. We propose Rubric-Calibrated Preferences (RCP), which combine both kinds of judgment. A listwise Bradley-Terry tournament orders each query's documents, and a rubric of yes/no criteria of increasing stringency provides an absolute standard. Item Response Theory (IRT), which scores test-takers based on their answers to common questions, then uses the shared criteria to put all queries' tournament scores on one scale. RCP's retrieval metric, RCP-nDCG, replaces nDCG's discrete labels with the resulting calibrated relevance probabilities. Against blind grades from 46 external annotators, calibration raises the correlation between a query's mean score and its mean human grade from 0.538 to 0.795. The probabilities rank a useful document above a non-useful one with probability 0.910 (AUC, chance 0.5), versus 0.651 for the benchmark labels. When the annotators' grades prefer one of two rerankers and exactly one metric agrees, that metric is RCP-nDCG in 72.4% of 185 comparisons (chance about 53%). On TREC-DL, RCP-nDCG sides with NIST assessors' grades on every reranker pair that these grades separate significantly. RCP-nDCG also resolves many of nDCG's ties and separates 1.9 times as many reranker pairs on NanoBEIR. Rubric calibration thus turns relative LLM judgments into dense relevance labels that are comparable across queries and agree with human judgment.

发表机构

  • Cohere

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑