LangBP:用于联合出价与定价的语言引导推理与行动
LangBP: Language-Guided Reasoning and Acting for Joint Bidding and Pricing
浏览论文内容
中文总结 AI 辅助
该研究针对联合出价与定价任务,提出LangBP分层框架,通过S-DT和EGPO解决现有语言引导方法的局限,在AuctionNet实验及电商平台部署中均取得良好效果。
中文摘要 AI 辅助
自动出价是在预算和关键绩效指标(KPI)约束下最大化转化价值的长期序贯决策问题。近期研究将该任务从单纯出价扩展至联合出价与定价,其中策略控制出价决策和定价修正。现有方法主要依赖数值轨迹建模,对活动上下文的解释和高级策略的表达支持有限。大语言模型(LLM)可凭借其推理能力补充这一范式,但现有语言引导方法存在两个局限:一是它们基于语言策略条件化动作,未对相应状态变化建模,难以区分策略理解错误与动作生成错误;二是不同指令可产生相似执行效果,导致跨效果的策略更新不平衡。我们提出LangBP,这是一个用于语言引导联合出价与定价的分层框架。LangBP的语义决策Transformer(S-DT)根据指令和轨迹历史预测目标状态,再通过逆动力学恢复联合动作。我们进一步提出执行分组策略优化(EGPO),它通过上下文-效果验证器(CEV)对候选效果评分,并平衡跨效果组的策略更新。在AuctionNet上的实验表明,LangBP优于强基线,在线A/B测试进一步证明其在大型电商平台实际部署中可带来业务收益。
英文摘要
Auto-bidding is a long-horizon sequential decision problem for maximizing conversion value under budget and key performance indicator (KPI) constraints. Recent work extends this task from bidding alone to joint bidding and pricing, where a policy controls bidding decisions and pricing corrections. Existing methods mainly rely on numerical trajectory modeling, which offers limited support for interpreting campaign context and expressing high-level strategies. Large language models (LLMs) can complement this paradigm with their reasoning capabilities. However, existing language-guided methods have two limitations. First, they condition actions on language strategies without modeling the corresponding state changes, making it difficult to distinguish errors in strategy understanding from errors in action generation. Second, different instructions can produce similar execution effects, leading to imbalanced policy updates across effects. We propose LangBP, a hierarchical framework for language-guided joint bidding and pricing. LangBP's Semantic Decision Transformer (S-DT) predicts target states from the instruction and the trajectory history, then recovers the joint action via inverse dynamics. We further propose Execution-Grouped Policy Optimization (EGPO), which scores candidate effects with a Context--Effect Verifier (CEV) and balances policy updates across effect groups. Experiments on AuctionNet show that LangBP outperforms strong baselines, and online A/B tests further demonstrate business gains in real-world deployment on a large-scale e-commerce platform.
发表机构
- Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。