arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2310.10696cs.LGcs.AI

面向流行度分布偏移的鲁棒协同过滤

Robust Collaborative Filtering to Popularity Distribution Shift

An Zhang, Wenchang Ma, Jingnan Zheng, Xiang Wang, Tat-seng Chua

首次发表 更新
浏览论文内容

中文总结 AI 辅助

针对协同过滤中流行度分布偏移问题,提出去偏策略PopGo,无需测试数据假设,量化并减少交互层面流行度捷径,在ID和OOD测试集上显著提升MF和LightGCN性能。

中文摘要 AI 辅助

在领先的协同过滤(CF)模型中,用户和物品的表征容易将训练数据中的流行度偏差作为捷径进行学习。流行度捷径技巧有利于分布内(ID)性能,但难以泛化到分布外(OOD)数据,即当测试数据的流行度分布相对于训练数据发生偏移时。为缩小这一差距,去偏策略试图评估捷径程度并从表征中减轻其影响。然而,存在两个缺陷:(1)在衡量捷径程度时,大多数策略仅使用单一方面的统计指标(即物品侧频率和用户侧频率),未能适应用户-物品对的组合程度;(2)在减轻捷径时,许多策略假设测试分布事先已知。这导致去偏表征质量低下。更糟糕的是,这些策略以牺牲ID性能为代价来获得OOD泛化能力。在本工作中,我们提出了一种简单而有效的去偏策略PopGo,它在不对测试数据做任何假设的情况下,量化并减少交互层面的流行度捷径。它首先学习一个捷径模型,该模型基于用户和物品的流行度表征给出用户-物品对的捷径程度。然后,通过用交互层面的捷径程度调整预测来训练CF模型。通过从因果和信息论两个角度审视PopGo,我们可以证明它为何能促使CF模型捕获关键的流行度无关特征,同时排除虚假的流行度相关模式。我们使用PopGo对四个基准数据集上的两个高性能CF模型(MF、LightGCN)进行去偏。在ID和OOD测试集上,PopGo均显著优于最先进的去偏策略(如DICE、MACR)。

英文摘要

In leading collaborative filtering (CF) models, representations of users and items are prone to learn popularity bias in the training data as shortcuts. The popularity shortcut tricks are good for in-distribution (ID) performance but poorly generalized to out-of-distribution (OOD) data, i.e., when popularity distribution of test data shifts w.r.t. the training one. To close the gap, debiasing strategies try to assess the shortcut degrees and mitigate them from the representations. However, there exist two deficiencies: (1) when measuring the shortcut degrees, most strategies only use statistical metrics on a single aspect (i.e., item frequency on item and user frequency on user aspect), failing to accommodate the compositional degree of a user-item pair; (2) when mitigating shortcuts, many strategies assume that the test distribution is known in advance. This results in low-quality debiased representations. Worse still, these strategies achieve OOD generalizability with a sacrifice on ID performance. In this work, we present a simple yet effective debiasing strategy, PopGo, which quantifies and reduces the interaction-wise popularity shortcut without any assumptions on the test data. It first learns a shortcut model, which yields a shortcut degree of a user-item pair based on their popularity representations. Then, it trains the CF model by adjusting the predictions with the interaction-wise shortcut degrees. By taking both causal- and information-theoretical looks at PopGo, we can justify why it encourages the CF model to capture the critical popularity-agnostic features while leaving the spurious popularity-relevant patterns out. We use PopGo to debias two high-performing CF models (MF, LightGCN) on four benchmark datasets. On both ID and OOD test sets, PopGo achieves significant gains over the state-of-the-art debiasing strategies (e.g., DICE, MACR).

发表机构

  • National University of Singapore(新加坡国立大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Institute of Dataspace, Hefei Comprehensive National Science Center(合肥综合性国家科学中心数据空间研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑