AI 中文总结
该研究将分组相对策略优化(GRPO)框架适配自动出价场景,经BAT等基准测试,其在点击量上优于基线,转化量为最佳或次佳方法。
AI 中文摘要
自动出价(autobidding)是现代在线广告系统的核心组件。在该组件中,广告商将序列出价决策委托给算法,这些算法必须在遵守预算有限、目标单次点击成本(CPC)等约束的同时最大化广告活动价值。解决自动出价问题的方法之一是将其表述为马尔可夫决策过程,并使用强化学习(RL)训练出价生成函数。标准的RL框架是演员-评论家(actor-critic),由生成动作的演员网络和估计这些动作价值的评论家网络组成。在我们的场景中,动作通常是出价或相关的 pacing 乘数,而价值是给定出价下拍卖的预期回报。然而,演员-评论家RL模型的交替训练会导致不稳定,且对噪声的鲁棒性降低。为解决这些问题,我们将分组相对策略优化(GRPO)框架适配到自动出价场景中。该框架是一种无评论家(critic-free)的策略梯度方法,最初开发用于大型语言模型(LLM)的后训练,其中真实目标未知。自动出价场景也具有这一特性,因为最优出价预先未知。此外,LLM领域的GRPO用于微调预训练模型,我们采用相同技术提升强启发式基线的性能。我们在BAT、iPinYou和AuctionNet基准上,将自动出价GRPO与演员-评论家模型、简单启发式方法及基于控制器的方法进行实证比较。大量实验表明,自动出价GRPO在点击量上始终优于所有基线,且在转化量上是最佳或第二优的方法。
英文摘要
Automated bidding (autobidding) is a core component of modern online advertising systems. Within this component, advertisers delegate sequential bid decisions to algorithms that must maximize campaign value while adhering to constraints such as a limited budget and a target cost-per-click (CPC). One of the approaches to resolve the autobidding problem is to formulate it as a Markov decision process and use reinforcement learning (RL) to train a bid generation function. The standard RL framework is actor-critic, which consists of an actor network that generates actions and a critic network that estimates the value of those actions. In our setting, the action is typically a bid or related pacing multiplier, and the value is the expected return from the auction given the bid. However, the alternating training of actor-critic RL models leads to instability and reduced robustness to noise. To address these issues, we adapt the Group Relative Policy Optimization (GRPO) framework to the autobidding setting. This framework is a \emph{critic-free} policy-gradient method originally developed for large language model post-training, where the ground-truth target is unknown. The autobidding setting shares this property, since the optimal bid is unknown in advance. Moreover, GRPO in the LLM domain is used to fine-tune the pre-trained model, and we use the same technique to enhance the performance of the strong heuristic baseline. We empirically compare Autobidding GRPO with actor-critic models, simple heuristics, and controller-based methods on the BAT, iPinYou, and AuctionNet benchmarks. Extensive experiments show that Autobidding GRPO consistently outperforms baselines in clicks and is the best or second-best method in conversion volume.