发表机构
Texas A&M University; University of New South Wales(德克萨斯农工大学; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种两阶段框架,通过角度相似性评分和仅微调偏置与层归一化组件,在不缓存历史的情况下关闭序列推荐中长短视角的性能差距,并在三个公共数据集上验证了有效性。
AI 中文摘要
序列推荐器通常在长用户历史上进行训练,以捕捉丰富的行为信号,但由于实时效率限制,使用训练长度序列进行服务往往不切实际。仅直接使用近期行为会导致严重的性能下降。为弥合这一差距,现有方法将用户历史压缩为持久的每用户状态,在推理时存储和检索这些状态;虽然有效,但它们带来了不小的基础设施开销,并且在冷启动场景中几乎没有补救措施。在本文中,我们通过经验识别出两个源于几何属性和数据集稀疏性的结构缺陷,并提出了一种新颖的两阶段框架来关闭长短视角性能差距。具体而言,在第一阶段,我们用角度相似性评分替代常用的点积,并利用修改后的softmax来抵消前缀位置偏差。在第二阶段,我们仅微调偏置和层归一化组件,这些组件对标准序列骨干是通用的,以进一步改进。两个阶段均由精心设计的学习目标指导。在两个代表性骨干上跨三个公共数据集的大量实验证明了我们提出框架的有效性。
英文摘要
Sequential recommenders are typically trained on long user histories to capture rich behavioral signals, yet serving with training-length sequences is often impractical due to real-time efficiency constraints. Directly using only recent behaviors leads to a severe performance drop. To bridge this gap, existing approaches compress user histories into persistent per-user states, storing and retrieving them at inference time; while effective, they impose non-trivial infrastructure overhead and offer little remedy in cold-start scenarios. In this paper, we empirically identify two structural flaws rooted in geometric properties and dataset sparsity, and propose a novel two-stage framework to close the long-short-view performance gap. Specifically, in the first stage, we replace the commonly used dot-product with angular similarity scoring and leverage a modified softmax to counter prefix position bias. In the second stage, we fine-tune only bias and LayerNorm components, which are universal to standard sequential backbones, for further improvement. Both stages are guided by carefully designed learning objectives. Extensive experiments on two representative backbones across three public datasets demonstrate the effectiveness of our proposed framework.
CommentsAccepted at CIKM 2026