AI 中文总结
研究针对离线 Top-N 推荐评估依赖不完整相关性信息的问题,提出基于大语言模型的框架,通过归纳用户偏好及判断候选项目相关性来扩展判断,提升评估稳健性并减轻流行度敏感失真。
AI 中文摘要
离线评估是比较 Top-N 推荐系统的标准方法,但依赖不完整的相关性信息。在多数基准数据集中,仅观察到一小部分用户-项目偏好,未判断项目常被视为不相关。这种缺失即负向的假设会使评估产生偏差。我们提出基于大语言模型的框架来扩展离线推荐评估的相关性判断。该方法利用大语言模型的两个互补作用,先归纳用户历史交互形成文本概况,再据此判断候选推荐项目的相关性。通过对多个推荐器排名靠前输出构建的候选集进行判断扩展,实验结果表明此方法能提升离线 Top-N 评估的稳健性并减轻稀疏反馈导致的流行度敏感失真。
英文摘要
Offline evaluation is the standard methodology for comparing top-N recommender systems, yet it relies on incomplete relevance information. In most benchmark datasets, only a small subset of user--item preferences is observed, and unjudged items are commonly treated as non-relevant. This missing-as-negative assumption can bias evaluation, penalize plausible recommendations with no recorded feedback, and favour algorithms that concentrate on popular or highly exposed items. We propose an LLM-based framework to expand relevance judgements for offline recommender evaluation. Our approach uses large language models in two complementary roles. First, a preference induction stage summarizes each user's historical interactions into a textual profile that captures their tastes and interests. Second, conditioned on this profile, an LLM acts as a relevance judge for candidate recommended items that lack observed labels in the original test data. To make this process tractable and evaluation-focused, we apply judgement expansion to a pooled candidate set built from the top-ranked outputs of multiple recommenders. The resulting enriched judgements provide additional relevance evidence for previously unobserved user--item pairs, enabling ranking metrics to be computed on a more complete basis. Experimental results show that this approach is a promising strategy for improving the robustness of offline top-N evaluation and mitigating the popularity-sensitive distortions caused by sparse feedback.