arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCOReD:用于推荐蒸馏的学生感知思维链优化

SCOReD: Student-Aware CoT Optimization for Recommendation Distillation

Haz Sameen Shahgir, Yufei Li, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Frank Shyu, Sandeep Pandey, Luke Simon, Yue Dong, Xi Liu

arXiv 2607.05734首次发表:更新:

发表机构

University of California Riverside; Meta AI(加州大学河滨分校; Meta AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对推荐蒸馏中原始教师轨迹不适用于任务的问题,提出SCOReD框架,先解析教师轨迹片段并评分,再动态选择编辑操作,使轨迹适应学生输出分布,训练能为学生模型提供清晰信号,提升指标并减少推理长度。

AI 中文摘要

推荐领域中的思维链蒸馏是强化学习训练的必要前提,但原始教师轨迹不适用于此任务。大型教师在处理推荐任务时推理不确定性异常高,反复检查答案却不修改;在此类轨迹上进行监督微调会产生从不修改初始猜测的冗长学生模型。此外,由于推荐领域的新颖性,教师的推理轨迹对小型学生语言模型来说分布严重不同。我们提出了用于推荐蒸馏的学生感知思维链优化(SCOReD),这是一个针对推荐量身定制的思维链优化框架,它首先将每个教师轨迹解析为类型化片段,并利用学生语言模型的注意力对每个片段的重要性进行评分。然后,SCOReD根据学生给出的编辑后的答案的输出长度和比较对数概率提升,动态地为每个片段选择编辑操作(保留/重写/融合/修剪)。因此,SCOReD在保留信息密集部分的同时修剪推理轨迹的冗余部分,并使原始教师轨迹适应学生的输出分布。在SCOReD优化的思维链上进行训练为学生模型提供了更清晰的学习信号,在归一化折损累计增益(NDCG)上比基线监督微调提高了1.56%,在召回率@5上提高了1.9%,同时推理长度减少了27.3%。

英文摘要

Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-suited to this task. Large teachers approach the recommendation task with unusually high reasoning uncertainty, repeatedly rechecking their answers without revising them; supervised fine-tuning on such traces produces verbose students that never revise their initial guess. Furthermore, due to the novelty of the recommendation domain, the teacher's reasoning traces are highly out-of-distribution for the small student LLM. We propose Student-Aware CoT Optimization for Recommendation Distillation (SCOReD), a CoT optimization framework tailored to recommendation that first parses each teacher trace into typed segments and uses the student LLM's attention to score the importance of each segment. Then SCOReD dynamically selects a per-segment edit (KEEP / REWRITE / FUSE / PRUNE) based on the output length and comparative log probability lift of the answer given the edit as per the student. Therefore, SCOReD prunes redundant sections of the reasoning trace while preserving information-dense sections and adapts raw teacher traces to the student's output distribution. Training on SCOReD-optimized CoTs provides a cleaner learning signal to the student model and improves over baseline SFT by 1.56% NDCG and 1.9% Recall@5, while reducing reasoning length by 27.3%.

Comments31 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑