稀疏且含噪反馈下生成式推荐器微调的指数奖励加权方法
Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback
浏览论文内容
中文总结 AI 辅助
针对稀疏含噪反馈下生成式推荐器的过优化问题,提出Exp-RSFT方法,通过指数奖励加权和温度λ平衡覆盖与噪声成本,在多数据集上验证其优于PPO、DPO,可提升排序性能且无需额外数据。
中文摘要 AI 辅助
在推荐系统中,用户仅与海量物品目录中的一小部分产生交互,产生的反馈既稀疏又含噪。这对训练后的生成式推荐器构成挑战:从记录的交互中训练的奖励模型往往无法泛化,而直接优化不完美的奖励会导致奖励过优化。我们提出指数奖励加权微调(Exp-RSFT),其中每个记录的交互按$\boldsymbol{\textit{exp}(r/\boldsymbol{\textit{λ}})}$加权,通过直接优化记录的奖励来避免上述问题,温度$\boldsymbol{\textit{λ}}$可正则化反馈噪声。我们从理论上证明,Exp-RSFT的次优性可分解为两类成本:日志策略局限性导致的覆盖成本,以及不完美反馈带来的噪声成本。温度$\boldsymbol{\textit{λ}}$可平衡这两类相互冲突的效应,在利用高奖励行为与抗噪声能力间实现最优权衡。在三个公开基准和一个大规模工业数据集上,我们验证了该理论预测:性能随$\boldsymbol{\textit{λ}}$呈倒U型变化,而PPO和DPO常过度优化不可靠的奖励模型,导致推荐质量下降。Exp-RSFT可持续提升排序性能,无需在线探索或偏好数据。
英文摘要
In recommendation systems, users interact with only a small fraction of a vast item catalog, producing feedback that is both sparse and noisy. This challenges post-training generative recommenders: reward models trained from logged interactions often fail to generalize, while directly optimizing imperfect rewards can lead to reward over-optimization. We propose Exponential reward-weighted fine-tuning (Exp-RSFT), where each logged interaction is weighted by $\exp(r/λ)$, avoids this failure by optimizing directly on the logged rewards, with the temperature $λ$ regularizing against their noise. We theoretically show that Exp-RSFT's suboptimality decomposes into two costs: a coverage cost arising from limitations of the logging policy and a noise cost from imperfect feedback. The temperature $λ$ balances these competing effects, yielding an optimal tradeoff between exploiting high-reward behavior and robustness to noise. Across three public benchmarks and a large-scale industrial dataset, we verify this theoretical prediction: performance follows an inverted-U trend as a function of $λ$, while PPO and DPO often over-optimize unreliable reward models and degrade recommendation quality. Exp-RSFT consistently improves ranking performance without requiring online exploration or preference data.