AI 中文总结
研究人员提出适用于多答案检索与排序评估的RDQ指标,该指标基于序数排序,兼顾返回项与排序,在POI数据集和TREC基准测试中展现出良好性能。
AI 中文摘要
我们提出了排名偏差质量(Rank-Deviation Quality, RDQ),这是一种适用于检索与排序系统的评估指标,可适配参考项数量不同的查询场景,从单个正确答案到多个有效结果均可覆盖。RDQ会针对有序参考列表(Ordered Reference List, ORL)对候选排序进行打分:每个检索到的参考项贡献其输出位置权重乘以排名偏差惩罚值,ORL之外的项则获得零分。应用特定参数可控制对错误排序的容忍度:较大的参数值更强调检索到有效参考项,较小的参数值则更重视匹配参考项的排序。输出位置权重可反映应用界面中的可见性,例如垂直列表或轮播。与需要绝对相关性等级的指标不同,RDQ基于序数排序运行,标注者可通过成对或列表式判断生成此类排序;与Kendall's tau等秩相关度量不同,RDQ同时考虑返回的项及其排序方式。在包含5000个查询的兴趣点(Point-of-Interest, POI)数据集及12个系统的实验中,RDQ在13种被评估的指标配置中具有最高的中位数经验power@100;在200个查询时,RDQ达到与自身全查询排序的平均tau≥0.8的一致性,而测试中最强的非RDQ配置RBP(0.9)需250个查询才能达到相同阈值。在TREC深度学习基准测试中,NDCG使用原生分级标签,RDQ使用从其衍生的序数层级,RDQ在n=25时达到可比的中位数power,而NDCG在n=100时更高。
英文摘要
We introduce Rank-Deviation Quality (RDQ), an evaluation metric for retrieval and ranking systems that adapts to queries with varying numbers of reference items, from a single correct answer to many valid results. RDQ scores a candidate ranking against an ordered reference list (ORL): each retrieved reference item contributes its output-position weight multiplied by a rank-deviation penalty, and items outside the ORL receive zero credit. Application-specific parameters control tolerance to misordering. Larger values emphasize retrieving valid reference items, whereas smaller values place more weight on matching their reference order. The output-position weights can reflect visibility in the application's interface, such as a vertical list or a carousel. Unlike metrics that require absolute relevance grades, RDQ operates on ordinal rankings, which annotators can produce through pairwise or listwise judgments. Unlike rank-correlation measures such as Kendall's tau, RDQ accounts for both which items are returned and how they are ordered. On a 5,000-query point-of-interest (POI) dataset with 12 systems, RDQ has the highest median empirical power@100 among the 13 evaluated metric configurations. It reaches mean tau >= 0.8 agreement with its own full-query ordering at 200 queries; RBP(0.9), the strongest tested non-RDQ configuration, reaches the same threshold at 250. On TREC Deep Learning benchmarks, where NDCG uses native graded labels and RDQ uses ordinal tiers derived from them, RDQ reaches comparable median power at n=25, while NDCG is higher at n=100.