arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HOBA:用于自适应在线广告的分层策略性投标代理

HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising

Ji Wu, Yunshan Peng, Wentao Bai, Yunke Bai, Wenzheng Shu, Jinan Pang, Yanxiang Zeng, Xialong Liu

arXiv 2607.24779首次发表:更新:

发表机构

Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线广告投标系统问题,提出 HOBA 分层强化学习框架,在三个时间尺度解耦相关环节,通过实验验证其优于现有基线,大规模在线部署实现目标成本提升 3.6%,证明该范式有效。

AI 中文摘要

在线广告投标系统通常部署多个离线训练的专家模型,但面临两个关键限制:缺乏对非平稳拍卖市场的在线适应性,以及依赖诸如投标界限和预算 pacing 约束等超参数的昂贵手动调整。我们提出了 HOBA,这是一个分层强化学习框架,它在三个时间尺度上解耦了战略推理、模型选择和投标执行。实验表明 HOBA 优于现有基线,在大规模在线部署中,HOBA 实现了目标成本增加 3.6%,证明了分层多智能体投标范式有效。

英文摘要

Online advertising bidding systems typically deploy multiple offline-trained expert models (e.g., PID controllers, model predictive control, offline RL policies) but face two critical limitations: lack of online adaptability to non-stationary auction markets, and reliance on costly manual tuning of hyperparameters such as bid bounds and budget pacing constraints. We propose HOBA (Hierarchical On-policy Bidding Agents), a hierarchical reinforcement learning framework that decouples strategic reasoning, model selection, and bid execution across three time scales. At the high level, a large language model infers hyperparameters from contextual signals through a Think-Act-Observe-Reflect loop with historical experience retrieval. At the mid level, a SARSA agent dynamically selects among expert models, incorporating causal adjustment to eliminate selection bias. At the low level, a dynamic expert pool (PID, MPC, IQL, Decision Transformer) executes bids under high-level constraints. This design confines online learning to discrete expert selection rather than continuous bid optimization, significantly reducing exploration risk while maintaining adaptability. Experiments on the AuctionNet benchmark and a large-scale A/B test demonstrate consistent improvements over state-of-the-art baselines. In a large-scale online deployment, HOBA delivered substantial business value, achieving a +3.6\% increase in target cost, proving the effectiveness of our hierarchical multi-agent bidding paradigm.

Comments10pages,accepted by KDD 2026 ads track

DOI:10.1145/3770855.3818435

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑