arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27953cs.IRcs.LG

LLM推荐重排序的召回上限

The Recall Ceiling of LLM Recommendation Reranking

Zhaohui Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示LLM推荐重排序的oracle评估高估性能,源于现实检索召回上限,提出RAEP协议以区分检索与重排序的改进优先级。

中文摘要 AI 辅助

一些基于LLM的推荐重排序器在oracle协议下进行评估,该协议保证真实物品存在于被评分的集合中,要么通过将其注入候选列表,要么通过将其与采样负例进行评分。在三个主要的Amazon数据集上,我们表明该协议将现实中的NDCG@10高估了92%至95%。原因在于召回上限:在三个领域的八个数据集中,现实检索在$K=100$时仅覆盖2%至19%的相关物品,这为任何封闭候选重排序器的top-$k$ NDCG施加了确定性上界。在留一评估下,$\u0395[\u006E\u0064\u0063\u0067@k] \u2264 \u0052\u0065\u0063\u0061\u006C\u006C@|W_π|$,其中$W_π$是重排序器的候选窗口。在现实检索下,在我们主要的Amazon数据集上,所测试的优化策略均未显著优于协同过滤基线。这些策略包括提示工程、在168倍参数范围内的模型扩展、序列模型、监督神经重排序器、LoRA微调、混合检索、分数感知提示以及LLM+CF融合。文本感知检索在一个数据集上增加了召回率,但并未改善端到端NDCG,而提供上游CF分数主要使LLM复现CF顺序。因此,我们提出了召回感知评估协议(RAEP):首先分类检索-召回机制,然后在上限允许有意义区分的条件下评估重排序。在本文测得的低召回机制中,改进检索比增加重排序器的复杂性更为重要;这种排序在生产系统中不一定成立,因为生产系统具有更高的召回率、更丰富的特征或在线反馈。

英文摘要

Some LLM-based recommendation rerankers are evaluated under an oracle protocol that guarantees the ground-truth item is present in the scored set, either by injecting it into the candidate list or by scoring it against sampled negatives. Across three primary Amazon datasets, we show that this protocol overestimates realistic NDCG@10 by 92--95%. The cause is a recall ceiling: realistic retrieval covers only 2--19% of relevant items at $K=100$ across eight datasets in three domains, imposing a deterministic upper bound on any closed-candidate reranker's top-$k$ NDCG. Under leave-one-out evaluation, $\mathbb{E}[\mathrm{NDCG}@k] \leq \mathrm{Recall}@|W_π|$, where $W_π$ is the reranker's candidate window. Under realistic retrieval, none of the tested optimisation strategies significantly improves over the collaborative-filtering baseline on our primary Amazon datasets. These strategies include prompt engineering, model scaling over a 168$\times$ parameter range, sequential models, supervised neural rerankers, LoRA fine-tuning, hybrid retrieval, score-aware prompting, and LLM+CF fusion. Text-aware retrieval increases recall on one dataset but does not improve end-to-end NDCG, while providing upstream CF scores mainly makes the LLM reproduce the CF order. We therefore propose the Recall-Aware Evaluation Protocol (RAEP): first classify the retrieval-recall regime, then evaluate reranking where the ceiling permits meaningful differentiation. In the low-recall regimes measured here, improving retrieval is more consequential than increasing reranker sophistication; this ordering need not hold in production systems with higher recall, richer features, or online feedback.

发表机构

  • University of Southern California(南加州大学)
  • Viterbi School of Engineering(维特比工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑