发表机构
Hokkaido University(北海道大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对基于Transformer的序列推荐器的最后一项依赖问题,通过分析注意力模块的残差主导性,揭示了其结构机制,为优化推荐器提供了依据。
AI 中文摘要
基于Transformer的序列推荐器采用因果自注意力机制,在推理时往往严重依赖最近一次交互,但这种行为如何在用于预测的表示中进行结构表达尚不清楚。我们将预测时诊断与完整注意力模块的基于范数的分析相结合。首先,我们证明SASRec风格模型表现出高度局部化的最后一项依赖。随后发现,尽管自注意力会聚合上下文信息,但残差加法会将完整模块的表示急剧转向相同位置的贡献,我们将此称为残差主导性。为探究该解释,我们采用推理时残差缩放作为受控诊断干预。改变残差强度会在结构混合与最后一项依赖之间引发单调权衡,而降低残差强度可恢复一部分最终位置遗漏,其中非最终位置的表示已能正确对真实物品进行排序。我们的结果提供了一种结构解释,将极端最后一项依赖与推理时的残差主导性关联起来。代码公开可用。
英文摘要
Transformer-based sequential recommenders with causal self-attention often rely heavily on the most recent interaction at inference time, but how this behavior is structurally expressed in the representation used for prediction remains unclear. We combine prediction-time diagnostics with norm-based analysis of the full attention block. First, we show that SASRec-style models exhibit highly localized last-item reliance. We then find that, although self-attention aggregates contextual information, residual addition sharply shifts the full-block representation toward same-position contributions, which we term residual dominance. To probe this interpretation, we use inference-time residual scaling as a controlled diagnostic intervention. Changing the residual strength induces a monotonic trade-off between structural mixing and last-item reliance, while reducing residual strength recovers a subset of final-position misses for which representations at non-final positions already rank the ground-truth item correctly. Our results provide a structural account linking extreme last-item reliance to residual dominance at inference time. The code is publicly available.
Comments11 pages, 6 figures. Accepted at the 20th ACM Conference on Recommender Systems (RecSys'26)