arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DMRL:用于广告推荐中技能优化的文档介导强化学习

DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation

Wei Zhang, Hongji Li, Song Sun, Peng Yu, Xue Yang, Lei Zhao, Peng Jiang

arXiv 2609.02170首次发表:更新:

发表机构

Shanghai Jiao Tong University; Kuaishou Technology(上海交通大学; 快手科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对广告推荐技能优化的奖励归因难题,提出DMRL框架,结合双相对策略优化与长期奖励预测器,在短视频广告平台上实现优于基线的性能。

AI 中文摘要

广告推荐需要持续调整复杂的系统参数,同时平衡商业收益与用户体验。近期研究引入了带有技能文档的大语言模型(LLMs)来辅助这一劳动密集型流程,但技能优化仍在很大程度上依赖提示驱动,缺乏将奖励归因于特定文档编辑的原则性机制。为解决这一局限,我们提出了文档介导强化学习(DMRL),这是一种将技能文档优化建模为结构化编辑动作序列的技能自进化框架。在DMRL中,上层智能体执行受控的文档编辑,而冻结的下层任务智能体通过A/B测试评估其效果。为解决信用分配和长期结果问题,我们引入了两个关键组件:(1)双相对策略优化(DRPO),一种用于鲁棒且风险感知的优势估计的后训练策略优化方法;(2)长期奖励预测器(LRP),通过解耦表示学习和交叉注意力迁移对总体异质性进行建模,从而估计长期结果。DMRL已部署在大规模短视频广告平台上,大量实证评估表明,DMRL在关键广告指标上的表现优于最先进的基线模型。

英文摘要

Advertising recommendation requires continuously tuning complex system parameters while balancing commercial returns and user experience. Recent work has introduced large language models (LLMs) with skill documents to assist this labor-intensive process, but skill optimization remains largely prompt-driven, lacking a principled mechanism to attribute rewards to specific document edits. To address this limitation, we propose Document-Mediated Reinforcement Learning (DMRL), a skill self-evolution framework that models skill document optimization as a sequence of structured editing actions. In DMRL, an upper-level agent performs controlled document edits, while a frozen lower-level task agent evaluates their effects through A/B testing. To address credit assignment and long-term outcomes, we introduce two key components: (1) Dual-Relative Policy Optimization (DRPO), a post-training policy optimization method for robust and risk-aware advantage estimation; and (2) Long-term Reward Predictor (LRP), which estimates long-term outcomes by modeling population heterogeneity with disentangled representation learning and cross-attention transfer. DMRL was deployed on a large-scale short-video ads platform and extensive empirical evaluation shows that DMRL outperforms state-of-the-art baselines across key advertising metrics

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑