LLM-OSDA:多轮大语言模型对话中原生广告的最优停止动态拍卖机制
LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations
浏览论文内容
中文总结 AI 辅助
针对多轮LLM对话中原生广告的时机与分配耦合问题,提出LLM-OSDA动态拍卖机制,结合贝尔曼最优停止等,在模拟实验中使净收益提升11%且用户留存相当。
中文摘要 AI 辅助
大语言模型原生广告将赞助内容直接嵌入模型生成的回复中,把销售单元从固定广告位转变为动态对话中的某一时刻。现有大语言模型广告拍卖机制主要在单次回复内运行,仅确定获胜者却未确定投放时机,该扩展并非易事:每个会话仅有一次原生插入机会,停止时间取决于出价,将时机与分配耦合,静态真实性论证不再适用。我们提出基于大语言模型的最优停止动态拍卖(LLM-OSDA),这是一种集成贝尔曼最优停止、获胜者分配和包络定价的动态单次点击成本拍卖。与出价无关的大语言模型层估计上下文点击质量并无缝渲染获胜广告,而出价仅进入承诺拍卖机制。在精确贝尔曼神谕下,预期折现点击分配对每个广告商的出价单调,对应的包络支付使出价真实在预期中弱占优。为实际部署,学习型停止网络(StopNet)近似贝尔曼动作值,其决策仅在停止边界附近与最优策略不同,并根据近似误差界定由此产生的激励损失。在模拟对话广告语料库上的实验表明,LLM-OSDA相比最强的固定时机基准将净收益提高11%,同时保持相当的用户留存率。代码位于此https URL。
英文摘要
LLM-native advertising embeds sponsored content directly into model-generated responses, shifting the unit of sale from a fixed slot to a moment within an evolving conversation. Existing LLM ad-auction mechanisms primarily operate within a single response, settling the winner but not the timing. The extension is nontrivial: with one native insertion opportunity per session, the stopping time depends on bids, coupling timing with allocation, so static truthfulness arguments no longer apply. We propose the LLM-based Optimal Stopping Dynamic Auction (LLM-OSDA), a dynamic cost-per-click auction that integrates Bellman optimal stopping, winner allocation, and envelope pricing. A bid-independent LLM layer estimates contextual click quality and seamlessly renders the winning ad, while bids enter only the committed auction mechanism. Under an exact Bellman oracle, the expected discounted-click allocation is monotone in each advertiser's bid, and the corresponding envelope payment makes truthful bidding weakly dominant in expectation. For practical deployment, a learned StopNet approximates the Bellman action values. We show that its decisions differ from the optimal policy only near the stopping boundary and bound the resulting incentive loss in terms of its approximation error. Experiments on a simulated conversational advertising corpus show that LLM-OSDA improves net revenue by 11 percent over the strongest fixed-timing baseline while maintaining comparable user retention. Code is at https://github.com/2025Fang2025/llm-osda.