发表机构
Gaoling School of Artificial Intelligence, Renmin University of China; Alibaba Group(中国人民大学高瓴人工智能学院; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对AIGB范式在自动出价中存在的问题,提出AIGB-R1框架,利用大语言模型推理能力,通过分层规划器与执行器模块、经验驱动循环、两阶段训练及新优化方法,经实验验证该框架在自动出价任务中的有效性。
AI 中文摘要
自动出价在在线广告中起着至关重要的作用,能自动调整出价以优化广告商的商业目标。新兴的人工智能生成出价(AIGB)范式广泛采用生成模型来优化出价策略,但存在离线数据集模式覆盖有限和任务状态理解不足的问题,阻碍了对最优策略的有效探索。大语言模型(LLMs)具有先验世界知识和推理能力,为克服这些限制提供了有前景的方法。然而,直接将LLMs应用于自动出价任务面临数值精度有限、幻觉和推理延迟等固有挑战。为解决这些限制,我们提出了AIGB-R1,这是一个分层的自我进化自动出价框架,旨在通过LLMs的推理能力增强人工智能生成出价,包括用于宏观策略规划的高级规划器模块和用于细粒度决策的低级执行器模块。在此基础上,我们设计了一个经验驱动的自我进化循环,从积累的经验中实现自主策略探索和优化。我们采用离线预训练和训练后对齐的两阶段管道,并构建了一个用于策略展开的交互式出价模拟环境。此外,我们提出了解耦组相对策略优化(D-GRPO),通过优势解耦实现端到端优化。在大规模公共数据集上的实验结果证明了AIGB-R1的有效性。
英文摘要
Auto-bidding plays an essential role in online advertising, automatically adjusting bids for advertisers to optimize their commercial goals. The emerging AI-Generated Bidding (AIGB) paradigm widely adopts generative modeling to optimize bidding strategies, yet suffers from the limited mode coverage of offline datasets and inadequate task-state understanding, hindering effective exploration of optimal strategies. Large Language Models (LLMs), with prior world knowledge and reasoning capabilities, offer a promising approach to overcome these limitations. However, directly applying LLMs to auto-bidding tasks faces inherent challenges in limited numerical precision, hallucinations, and inference latency. To address these limitations, we propose AIGB-R1, a hierarchical self-evolving auto-bidding framework aiming to enhance AI-Generated Bidding via LLMs' Reasoning capabilities, comprising a high-level Planner module for macro-level strategy planning and a low-level Executor module for fine-grained decision-making. Building upon this, we design an experience-driven self-evolving loop, enabling autonomous strategy exploration and optimization from accumulated experience. We adopt a two-stage pipeline of offline pre-training and post-training alignment, and build an interactive bidding simulation environment for strategy rollout. Furthermore, we propose Decoupled Group Relative Policy Optimization (D-GRPO) to achieve end-to-end optimization via advantage decoupling. Experimental results on a large-scale public dataset demonstrate the effectiveness of AIGB-R1.