arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31696cs.IRcs.LG

推荐系统离线评估中的数据预处理:综述

Data Processing for Offline Evaluation in Recommender Systems: a Survey

Alberto Carlo Maria Mancino, Angela Di Fazio, Danilo Danese, Matteo Attimonelli, Daniele Malitesta, Antonio Ferrara, Claudio Pomo, Tommaso Di Noia

AI总结:

本综述系统梳理推荐系统离线评估中的数据预处理实践,提出统一框架与分类法,并揭示数据转换与划分协议中的异质性问题。

AI中文摘要:

离线评估是推荐系统研究中的主导实验范式,它能够在历史交互数据上实现可重复且成本效益高的比较。然而,尽管推荐模型和评估方法已受到大量关注,但在模型训练之前的数据处理决策却较少受到审视。这些决策决定了推荐算法可获取的信息,并可能影响实验结果的可比性和可重复性。本综述对推荐系统离线评估的数据处理实践进行了系统性的跨领域特征刻画。我们考察了以数据为中心的流水线,从数据集选择和交互表示,到数据准备、多模态特征提取以及训练-验证-测试划分。我们的分析涵盖了多种推荐范式,包括协同过滤、序列推荐、基于会话的推荐、基于图的推荐、知识感知推荐、上下文感知推荐、多模态推荐、联邦推荐、跨域推荐、对比学习推荐以及基于大语言模型的推荐。除了回顾现有实践,我们引入了一个统一的框架和分类法,用于描述数据转换和特征提取策略,将数据准备与从多模态辅助信息中提取表示区分开来。我们的实证分析揭示了一个由狭窄的数据集级转换(尤其是基于支持度的过滤)主导的格局,而依赖表示转换的实践则相对少见。我们进一步发现,在辅助信息的准备和表示方式上存在显著的异质性,以及在数据划分协议规范上的不一致性,其中相似的标签可能掩盖不同的实验条件。

英文摘要:

Offline evaluation is the dominant experimental paradigm in recommender systems research, enabling reproducible and cost-effective comparisons on historical interaction data. Yet, while considerable attention has been devoted to recommendation models and evaluation methodologies, the data processing decisions that precede model training have received less scrutiny. These decisions determine the information available to recommendation algorithms and can affect the comparability and reproducibility of experimental results. This survey provides a systematic, cross-domain characterisation of data processing practices for the offline evaluation of recommender systems. We examine the data-centric pipeline, from dataset selection and interaction representation to data preparation, multimodal feature extraction, and train-validation-test splitting. Our analysis spans recommendation paradigms, including collaborative, sequential, session-based, graph-based, knowledge-aware, context-aware, multimodal, federated, cross-domain, contrastive-learning, and LLM-based recommendation. Beyond reviewing existing practices, we introduce a unified framework and taxonomy for describing data transformations and feature-extraction strategies, distinguishing data preparation from the extraction of representations from multimodal side information. Our empirical analysis reveals a landscape dominated by a narrow set of dataset-level transformations, particularly support-driven filtering, while representation-dependent transformations remain less common. We further identify substantial heterogeneity in how auxiliary information is prepared and represented, as well as inconsistencies in the specification of data splitting protocols, where similar labels may conceal different experimental conditions.

↑