从未来学习:用于序列推荐的特权自蒸馏
Learning from the Future: Privileged Self-Distillation for Sequential Recommendation
浏览论文内容
中文总结 AI 辅助
该研究提出PSD框架,利用训练时可用的未来交互作为特权信息,通过双注意力掩码实现自蒸馏,在不增加推理成本的前提下提升序列推荐性能。
中文摘要 AI 辅助
序列推荐器通常采用与推理对齐的因果(仅前缀)目标,基于独热下一个物品标签进行训练。这种监督虽与部署兼容,但几乎无法提供非目标物品间相对偏好的信息。然而,记录的交互序列包含额外监督源:目标之后的交互常能揭示用户意图的演变,使目标更易解释。我们将这些未来交互视为仅训练时可用的特权信息,仅在学习阶段提供,推理阶段不可用。这引出一个自然问题:未来交互能否提供更丰富的监督,同时保持训练与推理时预测对齐?我们提出特权自蒸馏(Privileged Self-Distillation, PSD)框架,该框架将学习时信息与推理时信息分离。PSD对同一骨干应用两种注意力掩码:未来感知视图生成基于过去和未来交互的特权教师分布,仅前缀视图生成部署所用的学生分布。蒸馏特权分布将未来交互转化为仅训练时的监督,而非推理时输入。由于两个视图共享同一骨干,教师的优势纯粹是信息性的,而非架构性的,无需单独预训练教师,且其监督可随学生演变调整。PSD还使用优势可达门,将蒸馏聚焦于观测前缀可能支持的教师信号,同时采用动量平均教师以获得稳定目标。该框架在单阶段端到端优化,部署模型与推理成本保持不变。在公共基准及多种骨干上的实验显示出一致的性能提升。
英文摘要
Sequential recommenders are commonly trained with one-hot next-item labels under a causal (prefix-only) objective aligned with inference. While deployment-compatible, this supervision offers little insight into relative preferences among non-target items. Yet logged interaction sequences contain an additional supervisory source: interactions following the target often reveal how user intent evolves, making the target easier to interpret. We treat these future interactions as training-only privileged information, available during learning but not at inference. This raises a natural question: can future interactions provide richer supervision while keeping training aligned with inference-time prediction? We propose Privileged Self-Distillation (PSD), a framework that separates learning-time information from inference-time information. PSD applies two attention masks to the same backbone: a future-aware view yields a privileged teacher distribution conditioned on past and future interactions, while a prefix-only view yields the student distribution used for deployment. Distilling the privileged distribution converts future interactions into training-only supervision rather than inference-time inputs. Since both views share a backbone, the teacher's advantage is purely informational, not architectural, removing the need for a separately pretrained teacher and letting its supervision adapt as the student evolves. PSD further uses an advantage-reachability gate to focus distillation on teacher signals likely supported by the observed prefix, along with a momentum-averaged teacher for stable targets. The framework is optimized end-to-end in a single stage, leaving the deployed model and inference cost unchanged. Experiments across public benchmarks and diverse backbones show consistent improvements.