arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2609.03313cs.IR

SelfDR:面向基于大语言模型的推荐的推理自蒸馏框架

SelfDR: Self-Distillation from Reasoning for LLM-Based Recommendation

  • Tsinghua University(清华大学)
  • Quan Cheng Laboratory(量子计算实验室)
  • DCST, Tsinghua University(清华大学计算机科学与技术系)
  • Meituan(美团)

机构由 AI 辅助整理,请以论文原文为准。

Chumeng Jiang, Jiayin Wang, Xinjie Lin, Zhiqiang Guo, Hengliang Luo, Min Zhang

AI总结:

该研究针对基于LLM的推荐中推理轨迹生成成本高的问题,提出SelfDR框架,通过同一LLM的自蒸馏提升推荐效果与效率,在三个公开数据集上验证了其有效性。

AI中文摘要:

大语言模型(LLMs)近期已成为强大的推荐系统骨干模型。为更好地激发其能力,推理被广泛引入以帮助LLMs解读丰富的文本信号并提升推荐准确率。然而,显式生成中间推理轨迹往往会产生大量计算成本,这限制了其在实际推荐系统中的部署。为应对这一挑战,我们提出了SelfDR,一种面向基于LLM的推荐的推理自蒸馏框架。SelfDR蒸馏LLM自身经推理增强的预测结果,直接生成推荐,在保持推理效率的同时提升推荐效果。框架的所有组件均基于同一基础LLM构建,不依赖任何外部模型。具体而言,教师推荐器通过以下方式构建:以下游性能为奖励训练推理器,使其生成针对性的推理依据,后续将这些依据融入教师的输入。随后,具有相同底层模型的用于直接推荐的学生推荐器,通过带有动态加权策略的自蒸馏向教师学习。在三个公开数据集上进行的大量实验验证了SelfDR的有效性、合理性与效率。代码可在该https URL获取。

英文摘要:

Large Language Models (LLMs) have recently emerged as powerful backbones for recommendation. To better elicit their capabilities, reasoning has been widely incorporated to help LLMs interpret rich textual signals and improve recommendation accuracy. However, explicitly generating intermediate reasoning traces often incurs substantial computational costs, which limits practical deployment in real-world recommender systems. To address this challenge, we propose SelfDR, a Self-Distillation from Reasoning framework for LLM-based Recommendation. SelfDR distills an LLM's own reasoning-enhanced predictions to produce recommendations directly, improving recommendation effectiveness while maintaining inference efficiency. All components in the framework are built on the same base LLM, without relying on any external models. Specifically, the teacher recommender is constructed by training a reasoner with downstream performance as the reward, enabling it to generate targeted rationales that are later incorporated into the teacher's input. A student recommender for direct recommendation, with the same underlying model, then learns from the teacher through self-distillation with a dynamic weighting strategy. Extensive experiments on three public datasets validate the effectiveness, rationality, and efficiency of SelfDR. Codes are available at https://github.com/JiangDeccc/SelfDistillation.

补充信息

↑