arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28065cs.AI

基于离线模型强化学习的激励式广告激励分配学习

Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning

Zilin Zhao, Han Yang, Tianpei Yang, Fangsheng Huang, Yanfei Cui, Kan Peng, Yi Li, Yiming Zong, Hao Zhang, Yinsong Xue

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对激励式广告的序列激励分配问题,提出离线模型强化学习框架,实验显示MB-IQL可显著提升人均净利润,验证了该框架的有效性。

中文摘要 AI 辅助

完成广告观看即可获得5美分奖金!在激励式广告中,平台在观测到下游广告收入前就向用户承诺奖金,以鼓励用户点击并完成广告。平台必须平衡预先承诺的激励与后续实现的收入:激励不足会丧失盈利机会,而激励过度则会降低净利润。由于当前的激励措施还会影响用户预期和未来参与度,激励分配是一个具有延迟收入、成本敏感性和留存效应的序列决策问题。现有研究未针对该场景研究决策算法:自动竞价假设存在可用广告机会,定向推广则在广告盈利流程外优化激励。我们将该问题建模为马尔可夫决策过程(MDP),并开发了用于成本可控序列激励分配的离线模型强化学习(Offline Model-Based RL)框架。该框架学习用户反馈和广告收入的世界模型,随后执行保守策略优化。独立的反事实评分器在留存日志上评估每个学习到的策略,支持无需高昂在线曝光的发布前选择。对大规模工业数据的实验及在线A/B测试显示,该评分器提供了稳定的离线信号。从因果推理到离线RL再到Offline-MBRL的部署路径进一步验证了该框架:MB-IQL相比TD3+BC将人均净利润提升7.96%,而退化为普通IQL时则降低6.56%(两者p值均小于0.0001)。

英文摘要

Complete your ad view and grab a 5-cent bonus! In incentivized advertising, a platform promises users a bonus before observing downstream ad revenue, encouraging them to click and complete ads. It must balance the incentive promised in advance against the revenue realized afterward: insufficient incentives forfeit monetization opportunities, whereas excessive incentives reduce net profit. Because current incentives may also shape user expectations and future engagement, incentive allocation is a sequential decision problem with delayed revenue, cost sensitivity, and carryover effects. Existing work has not studied decision-making algorithms for this setting. Auto-bidding assumes available ad opportunities, while targeted promotion optimizes incentives outside the ad monetization pipeline. We formulate the problem as an MDP and develop an offline model-based RL framework for cost-controllable sequential incentive allocation. It learns a world model of user feedback and ad revenue, then performs conservative policy optimization. An independent counterfactual scorer evaluates each learned policy on held-out logs, enabling pre-launch selection without costly online exposure. Experiments on large-scale industrial data and online A/B tests show that the scorer provides a stable offline signal. The deployment path from causal inference to offline RL and then Offline-MBRL further validates the framework: MB-IQL improves per-user net profit by 7.96\% over TD3+BC, whereas reverting to plain IQL reduces it by 6.56\% (both \(p<0.0001\)).

发表机构

  • State Key Laboratory of Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
  • ByteDance(字节跳动)
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑