arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过多阶段后训练实现推荐系统基础模型的渐进式对齐

Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training

Oseong Choi, Hoeinn Kim, Jihoon Lee, Byungsoo Kang, Taeyeong Jang

arXiv 2608.06792首次发表:更新:

发表机构

NAVER WEBTOON(NAVER WEBTOON)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出LP-FFT-RFT三阶段渐进式后训练框架,将推荐系统基础模型的下游适配与业务指标对齐分离,实验证实其能提升推荐质量。

AI 中文摘要

推荐系统基础模型(Foundation Model, FM)在对用户长序列行为进行建模方面展现出强大能力。在实际应用中,单个预训练基础模型通常通过监督微调(Supervised Fine-Tuning, SFT)适配到多样化的下游服务场景。然而,针对点击、点赞等任务特定目标进行优化,并不一定能使服务策略与决定推荐质量的业务指标保持一致。本文提出一种三阶段渐进式后训练框架,明确将下游适配与业务指标对齐分离。适配阶段分解为线性探测(Linear Probing, LP)和全量微调(Full Fine-Tuning, FFT):LP首先在冻结的预训练表示空间内稳定随机初始化的下游头,随后FFT联合优化全模型以适配目标任务。在该稳定策略基础上,强化微调(Reinforcement Fine-Tuning, RFT)利用学习到的奖励模型使模型对齐实际业务目标。与直接在稀疏业务目标上优化服务策略不同,本文在密集隐式反馈上训练策略,仅将业务指标监督用于奖励建模。离线实验表明,渐进式LP-FFT-RFT框架优于单阶段方案,且基于奖励的对齐比直接使用奖励模型自身排序能生成更强的服务策略。大规模在线A/B测试进一步显示,所提框架相比传统非基础模型基线提升了生产环境的推荐质量,参考实现可在指定URL获取。

英文摘要

Foundation model(FM) for recommendation has shown strong ability to model long-horizon sequential user behavior. In practice, a single pretrained foundation model is often adapted to diverse downstream serving surfaces through Supervised Fine-Tuning(SFT). However, optimizing task-specific objectives such as clicks or likes does not necessarily align the serving policy with the business metrics that determine recommendation quality. We propose a three-phase progressive post-training framework that explicitly separates downstream adaptation from business-metric alignment. The adaptation stage is decomposed into Linear Probing(LP) and Full Fine-Tuning(FFT): LP first stabilizes randomly initialized downstream heads within a frozen pretrained representation space, and FFT then jointly specializes the full model for the target task. On top of this stabilized policy, Reinforcement Fine-Tuning(RFT) aligns the model with practical business objectives using a learned reward model. Rather than directly optimizing the serving policy on sparse business targets, we train the policy on dense implicit feedback and use business-metric supervision only for reward modeling. Offline experiments show that the progressive LP-FFT-RFT framework outperforms single-phase alternatives, and that reward-based alignment yields a stronger serving policy than directly using the reward model itself for ranking. Large-scale online A/B tests further show that the proposed framework improves production recommendation quality over a conventional non-foundation baseline. A reference implementation is available at https://github.com/webtoon/rec-fm-progressive-alignment

Comments9 pages, 3 figures. Accepted to the 20th ACM Conference on Recommender Systems (RecSys '26), Industry Track

DOI:10.1145/3773078.3831870

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑