面向全局层级约束的激励式广告生成优化
Generative Optimization for Incentivized Advertising with Global Level Constraints
浏览论文内容
中文总结 AI 辅助
针对全局约束下激励式广告的优化难题,提出GOAL框架与SCPO方法,在降低ROI违规率的同时提升长期收入与用户留存。
中文摘要 AI 辅助
激励式广告通过分配货币或虚拟奖励驱动用户参与,其核心挑战是在严格全局约束下优化连续激励幅度。该问题因高频交互、延迟反馈及非马尔可夫用户动态(如疲劳)而复杂化,现有增量建模与约束强化学习方法的有效性受限。为应对这些挑战,我们提出GOAL——一种感知约束的生成框架,将激励分配建模为条件序列生成问题。GOAL基于用户历史与系统级全局压力直接生成激励幅度,并集成分层因果状态编码器以捕捉局部行为动态与长程依赖。为实现灵活的约束控制,我们引入Safe Constrained Policy Optimization(SCPO),该方法学习单个生成策略,可在一系列投资回报率(ROI)约束下泛化,无需重新训练。在大规模真实数据与合成疲劳感知环境上的实验表明,与强基线相比,GOAL在提升长期收入与用户留存的同时,大幅降低了ROI违规率。
英文摘要
Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous incentive magnitudes under strict global constraints. This problem is complicated by high-frequency interactions, delayed feedback, and non-Markovian user dynamics such as fatigue, which limit the effectiveness of existing uplift modeling and constrained reinforcement learning approaches. To address these challenges, we propose GOAL, a constraint-aware generative framework that formulates incentive allocation as a conditional sequence generation problem. GOAL directly generates incentive magnitudes conditioned on user histories and system-level global pressure, and integrates a hierarchical causal state encoder to capture both local behavioral dynamics and long-range dependencies. To enable flexible constraint control, we introduce \textbf{S}afe \textbf{C}onstrained \textbf{P}olicy \textbf{O}ptimization (SCPO), which learns a single generative policy that generalizes across a spectrum of ROI constraints without retraining. Experiments on large-scale real-world data and a synthetic fatigue-aware environment show that GOAL improves long-term revenue and user retention while substantially reducing ROI violation rates compared to strong baselines.
发表机构
- University of Electronic Science and Technology of China(电子科技大学)
- Kuaishou Technology(快手科技)
机构由 AI 辅助整理,请以论文原文为准。