AI 中文总结
该研究通过两种无需高阶序列表示的简单探针,发现广泛使用的序列推荐基准不适于衡量高阶序列建模的增益,为相关基准有效性提供了具体测试方法。
AI 中文摘要
序列推荐器越来越多地采用语言模型架构来捕捉复杂的、依赖上下文的交互,但目前尚不清楚广泛使用的基准是否真的需要这种建模能力。我们使用两种简单的、基于近期加权的成对探针来研究这个问题,这些探针无需学习高阶序列表示:序列规则(SeqRules)和我们的概率协作转移模型(PCTM)。采用eSASRec的评估协议,在三个亚马逊数据集上,至少有一种探针超过我们的eSASRec复现结果15-38%,在MovieLens-1M数据集上超过4.4%,但在MovieLens-20M数据集上落后27.3%;在其余四个数据集上,至少有一种探针也超过我们的带采样softmax的SASRec复现结果9-28%,这表明这些广泛使用的基准并不适合衡量高阶序列建模带来的增益。更广泛地说,将基于Transformer的模型与强大的近期加权成对探针进行比较,为基准能否有意义地衡量高阶序列建模的增益提供了具体的测试方法。
英文摘要
Sequential recommenders increasingly use language-model architectures designed to capture complex, context-dependent interactions. Yet it remains unclear whether widely used benchmarks actually require this modelling capacity. We investigate this question using two simple, recency-weighted pairwise probes that do not learn higher-order sequence representations: Sequential Rules (SeqRules) and our Probabilistic Collaborative Transition Model (PCTM). Using the evaluation protocol of eSASRec, at least one probe exceeds our eSASRec reproduction by 15-38% on three Amazon datasets and by 4.4% on MovieLens-1M, but trails it by 27.3% on MovieLens-20M. On the four remaining datasets, at least one probe also outperforms our sampled-softmax SASRec reproduction by 9-28%, suggesting that these widely used benchmarks are poorly suited to measuring gains from higher-order sequence modelling. More broadly, comparing Transformer-based models against strong recency-weighted pairwise probes provides a concrete test of whether a benchmark can meaningfully measure gains from higher-order sequence modelling.
CommentsAccepted at the 20th ACM Conference on Recommender Systems (RecSys 2026)