arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvoReason:用于生成式推荐中潜在推理的自演化推理原语引导在线策略蒸馏

EvoReason: Self-Evolving Reasoning Primitive-Guided On-Policy Distillation for Latent Reasoning in Generative Recommendation

Zhuang Zhuang, Zhipeng Wei, Rongfeng Guo, Shijie Li, Peng Zhao, Jie Chen, Fei Pan

arXiv 2607.29010首次发表:更新:

AI 中文总结

针对生成式推荐中潜在推理的原始CoT监督效果不佳问题,提出自演化框架EvoReason,通过推理原语引导的在线策略蒸馏实现显式推理监督与学生潜在推理空间的自适应对齐,提升推理迁移效果。

AI 中文摘要

生成式推荐得益于推理增强的推理过程,而潜在推理通过将中间推理过程编码为紧凑的连续表示,为延迟敏感型部署提供了高效范式。尽管其效率较高,但现有潜在推理方法通常依赖于将原始思维链(CoT)轨迹直接蒸馏为潜在表示,假设文本推理轨迹能提供充足监督。然而,推荐推理轨迹包含多样的推理过程,存在冗余表达和不稳定的推理路径,使得原始CoT监督对学习可迁移的潜在推理表示效果不佳。为应对这一挑战,我们提出EvoReason,一种自演化潜在推理框架,通过原语引导的在线策略蒸馏自适应地将显式推理监督与学生的潜在推理空间对齐。首先,EvoReason从高质量的智能体推荐轨迹中提取可复用的推理原语,每个原语捕捉一个关键推理行为,作为结构化教师推理的伪工具。然后,基于这些原语,我们为教师配备原语感知推理能力,使其能生成冗余减少、一致性提升的结构化CoT监督。最后,在潜在推理优化过程中,EvoReason引入自演化在线策略蒸馏机制,其中原语引导的推理过程根据学生的潜在推理结果演化。通过这种闭环协同演化,策略更新不断改进潜在推理行为,并根据生成的潜在推理结果进行优化,实现逐步更好对齐的CoT监督和更有效的推理迁移。

英文摘要

Generative recommendation benefits from reasoning-enhanced inference, and latent reasoning offers an efficient paradigm by encoding intermediate reasoning processes into compact continuous representations for latency-sensitive deployment. Despite its efficiency, existing latent reasoning approaches typically rely on directly distilling raw chain-of-thought (CoT) trajectories into latent representations, assuming that textual reasoning traces provide sufficient supervision. However, recommendation reasoning trajectories contain diverse reasoning processes with redundant expressions and unstable reasoning paths, making raw CoT supervision suboptimal for learning transferable latent reasoning representations. To address this challenge, we propose EvoReason, a self-evolving latent reasoning framework that adaptively aligns explicit reasoning supervision with the student's latent reasoning space through primitive-guided on-policy distillation. First, EvoReason extracts reusable reasoning primitives from high-quality agentic recommendation trajectories, where each primitive captures an essential reasoning behavior and serves as a pseudo-tool for structured teacher reasoning. Then, based on these primitives, we equip the teacher with primitive-aware reasoning capabilities, enabling it to generate structured CoT supervision with reduced redundancy and improved consistency. Finally, during latent reasoning optimization, EvoReason introduces a self-evolving on-policy distillation mechanism, where the primitive-guided reasoning process evolves according to the student's latent reasoning outcomes. Through this closed-loop co-evolution, policy updates continuously improve latent reasoning behaviors is refined according to the resulting latent reasoning outcomes, enabling progressively better-aligned CoT supervision and more effective reasoning transfer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑