AI 中文总结
研究推荐系统离线评估设计对比较有效性的影响,通过改变数据过滤阈值等关键因素评估一组推荐模型,测量不同评估设置下模型排名相关性,发现稀疏评估有效性因数据集和目标而异,无统一最佳设计。
AI 中文摘要
对历史交互日志进行离线评估是推荐系统最常用的评估方法。然而,此类评估依赖于稀疏、不完整或有偏差的数据,这引发了人们对常用评估设置能否可靠反映真实用户偏好的担忧。在这项工作中,我们研究离线评估设计选择如何影响推荐系统比较的有效性。我们在多个评估设置下评估了一组推荐模型,这些设置改变了诸如数据过滤阈值和候选集构建等关键因素。为评估这些配置的有效性,我们测量了从稀疏交互数据上的传统训练 - 测试分割获得的模型排名与基于密集真实用户反馈的评估排名之间的相关性。我们将这种一致性作为它们相对于真实用户偏好有效性的指标。我们的结果表明,稀疏评估的有效性取决于数据集和特定的密集评估目标,并且不存在统一的最佳离线评估设计。
英文摘要
Offline evaluation on historical interaction logs is the most common evaluation methodology for recommender systems. However, such evaluations depend on sparse, incomplete, or biased data, which raises concerns about whether commonly used evaluation setups reliably reflect true user preferences. In this work, we study how offline evaluation design choices affect the validity of recommender system comparisons. We evaluate a set of recommendation models across several evaluation setups that vary key factors such as data filtering thresholds and candidate set construction. To assess the validity of these configurations, we measure the correlation between model rankings obtained from conventional train-test splits on sparse interaction data and rankings from evaluations based on dense ground-truth user feedback. We use this agreement as an indication of their validity with respect to true user preferences. Our results show that the validity of sparse evaluation depends on the dataset and the specific dense evaluation targets, and that there is no uniformly best offline evaluation design.
Comments9 pages, Accepted at ACM RecSys 2026 Main Track