arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21550cs.LGcs.AI

OneBid:面向多样化oCPX广告场景的统一自动出价基础模型

OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

  • Kuaishou Technology(快手科技)

机构由 AI 辅助整理,请以论文原文为准。

Yewen Li, Peng Jiang, Yitian Li, Pengfei Lv, Xialong Liu, Peng Jiang, Qingpeng Cai

AI总结:

OneBid提出统一自动出价基础模型,通过双信号条件化、序列级MoE和CROP离线优化,在快手oCPX广告中实现整体+2.2%、ROAS场景+13.1%的ADVV提升。

AI中文摘要:

自动出价是计算广告的核心,其策略必须在经济约束下最大化广告主的转化价值。自动出价已从基于规则的控制器演变为强化学习和生成式方法(如决策Transformer,DT)。然而,这些方法日益与主流的优化每行动成本(oCPX)范式不匹配,该范式涵盖异构场景(如注册、购买),每个场景由独立模型服务,导致流水线碎片化且跨场景建模探索不足。受LLM等基础模型的启发,将这些oCPX场景统一到一个模型中面临三个挑战:多目标控制、严格延迟下的可扩展容量,以及安全的离线策略改进。我们提出OneBid,一个统一的自动出价基础模型,它从异构oCPX日志中学习可复用主干,并通过离线后训练将其适配到特定场景部署。基于DT,OneBid将单一返回目标(Return-to-Go)条件扩展为两个原子信号:用于转化价值的返回目标和用于成本比率的成本目标(Cost-to-Go),并在下一动作预测上加入价值感知正则化。为吸收分布异构性,我们设计了序列级混合专家(MoE)架构,其中共享专家编码跨场景知识,稀疏路由专家以低延迟捕获场景特定模式,从而随模型大小和数据实现一致扩展。在后训练阶段,我们通过评论家引导的相对离线策略优化(CROP)将主干与场景偏好对齐:学习到的评论家对候选动作进行组相对评分,避免GRPO风格微调的不安全在线探索,同时约束策略偏移以降低分布外(OOD)风险。通过在线A/B测试验证并已在快手全面部署,OneBid在oCPX广告上整体提升+2.2%的ADVV,在ROAS场景中峰值提升+13.1%。

英文摘要:

Auto-bidding is central to computational advertising, where strategies must maximize advertisers' conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevailing optimized cost-per-X (oCPX) paradigm, which spans heterogeneous scenarios (e.g., registration, purchase), each served by a separate model, leading to fragmented pipelines and underexploring cross-scenario modeling. Inspired by foundation models like LLMs, unifying these oCPX scenarios into one model raises three challenges: multi-objective control, scalable capacity under strict latency, and safe offline policy improvement. We present OneBid, a unified auto-bidding foundation model that learns a reusable backbone from heterogeneous oCPX logs and adapts it to scenario-specific deployments via offline post-training. Building on DT, OneBid extends single Return-to-Go conditioning to two atomic signals, Return-to-Go for conversion value and Cost-to-Go for cost ratio, plus value-aware regularization on next-action prediction. To absorb distributional heterogeneity, we design a sequence-level Mixture-of-Experts architecture, where shared experts encode cross-scenario knowledge and sparsely-routed experts capture scenario-specific patterns at low latency, yielding consistent scaling with model size and data. During post-training, we align the backbone with scenario preferences via Critic-guided Relative Offline Policy optimization (CROP): a learned critic scores candidate actions group-relatively, avoiding the unsafe online exploration of GRPO-style fine-tuning while constraining policy shift to reduce OOD risk. Validated via online A/B tests and fully deployed at Kuaishou, OneBid delivers an overall +2.2% ADVV gain on oCPX Ads, peaking at +13.1% in the ROAS scenario.

↑