arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推荐系统评估中的训练种子与模型选择稳定性

Training seeds and model-selection stability in recommender-system evaluation

Juan Manuel Rodriguez, Oleg Lesota, Antonela Tommasel

arXiv 2609.02499首次发表:更新:

发表机构

Aalborg University; Johannes Kepler University Linz; ISISTAN, CONICET-UNCPBA(奥尔堡大学; 约翰开普勒林茨大学; ISISTAN、CONICET-UNCPBA)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过固定数据划分、调整训练种子分析其对推荐系统评估的影响,发现种子变异可检测且会影响评估稳定性,建议将训练种子纳入评估协议。

AI 中文摘要

推荐系统实验通常依赖单一随机训练种子,假设运行间随机性对评估结论影响有限。但该假设存在风险,因为训练种子可能影响多种算法相关机制,包括参数初始化、小批量排序、Dropout、掩码、潜在采样及训练时负采样。本研究固定数据划分,在超参数配置间改变训练种子,从三个层面分析种子效应:用户级指标敏感性、基于验证的模型选择以及推荐列表一致性。结果显示种子变异通常可被检测到,其影响取决于配置是否明确区分、验证结果能否迁移至测试,以及相似分数是否对应相似的Top-k列表。研究发现报告单种子结果可能夸大推荐系统评估的稳定性,训练种子应被视为评估协议的一部分,而非偶然的实现噪声。

英文摘要

Recommender-system experiments often rely on a single random training seed, assuming that run-to-run stochasticity has limited impact on evaluation conclusions. This assumption is risky, as a training seed may influence several algorithm-dependent mechanisms, including parameter initialization, mini-batch ordering, dropout, masking, latent sampling, and training-time negative sampling. We examine this assumption by fixing the data partition and varying the training seed across hyperparameter configurations. We analyze seed effects at three levels: user-level metric sensitivity, validation-based model selection and recommendation-list agreement. Results show that seed variation is often detectable. Its impact depends on whether configurations are clearly separated, whether validation results transfer to test, and whether similar scores lead to similar top-$k$ lists. Findings suggest that reporting single-seed results can overstate the stability of recommender system evaluation, and that training seeds should be treated as part of the evaluation protocol rather than as incidental implementation noise.

CommentsAccepted RecSys 2026 (Research&Practice Notes)

DOI:10.1145/3773078.3841289

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑