发表机构
Capital One(第一资本)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对顺序推荐系统缺乏可解释性基准的问题,提出双模型掩码度量,系统评估十种XAI方法,发现梯度类方法最忠实,注意力权重需梯度加权才可靠。
AI 中文摘要
顺序推荐系统是现代个性化的核心,利用用户的历史交互序列来驱动下一步决策。深度学习模型,尤其是基于CNN和Transformer的架构,已被证明在捕捉这些历史中的时间依赖性方面非常有效。为了透明度和信任,理解哪些过去的交互驱动了给定的推荐变得越来越重要——无论是对于审计模型行为的开发者,还是对于寻求理由的用户而言。然而,赋予这些模型预测能力的非线性特性也使它们成为黑盒,使得难以将决策归因于特定的交互。尽管存在基于梯度、基于扰动和基于注意力的可解释性方法,但针对顺序推荐的忠实性系统性基准仍然缺失。我们通过引入一种双模型掩码度量来填补这一空白,其中一个模型提供每个时间步的归因分数,另一个单独训练的、对掩码鲁棒的探针模型测量由此导致的预测概率变化。使用该度量,我们在KuaiRand和MovieLens数据集上,对CNN、Transformer、SASRec和BERT4Rec骨干网络上的十种XAI方法进行了基准测试,并辅以时间归因模式、物品流行度混杂因素以及对输入损坏的鲁棒性分析。我们的主要发现是:(1)基于梯度的方法,特别是GradientSHAP和Integrated Gradients,产生了最忠实和鲁棒的归因;(2)原始注意力权重不可靠,但梯度加权注意力在较短序列上恢复了忠实性,而在较长序列上,随着softmax注意力概率收敛于均匀重要性分数,该方法识别信息性交互的能力下降;(3)忠实方法中的时间归因模式反映了真实的任务结构,而非近因或流行度偏差。
英文摘要
Sequential RecSys are central to modern personalization, exploiting user's historical interaction sequences to drive next-step decisions. Deep learning models, particularly CNN and Transformer-based architectures, have proven highly effective at capturing temporal dependencies in these histories. For transparency and trust, understanding which past interactions drive a given recommendation is increasingly important --- both for developers auditing model behavior and for users seeking a rationale. However, the non-linearities that give these models their predictive power also render them black boxes, making it difficult to attribute decisions to specific interactions. While gradient-based, perturbation-based, and attention-based explainability methods exist, a systematic benchmark of their faithfulness for sequential recommendation is missing. We address this gap by introducing a dual-model masking metric in which one model supplies per-timestep attribution scores and a separately trained, masking-robust probe measures the resulting change in predicted probability. Using this metric, we benchmark ten XAI methods across CNN, Transformer, SASRec, and BERT4Rec backbones on KuaiRand and MovieLens, complemented by analyses of temporal attribution patterns, item popularity confounding, and robustness to input corruption. Our key findings are: (1) gradient-based methods, particularly GradientSHAP and Integrated Gradients, yield the most faithful and robust attributions; (2) raw attention weights are unreliable, but gradient-weighted attention restores faithfulness on shorter sequences, with degradation on longer horizons as softmax attention probabilities converge toward uniform importance scores, diminishing the method's ability to identify informative interactions; and (3) temporal attribution patterns in faithful methods reflect genuine task structure rather than recency or popularity bias.
CommentsPresented at CARS@RecSys'26, 11 pages, 6 figures