在面对平滑对手的双方面贸易中的利润最大化
Profit Maximization in Bilateral Trade against a Smooth Adversary
- Google Research(谷歌研究)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究在在线学习框架下,面对平滑对手生成的估值,设计出保证$\tilde{O}(\sqrt{T})$ regret的算法,拓展了快速收敛的适用范围,填补了经济问题的regret景观中的空白。
AI中文摘要:
双方面贸易模型了中介两个战略代理体的任务,即卖家和买家希望交易商品。我们从在线学习框架中利润最大化经纪人的视角研究此问题,其中代理体的估值由平滑对手生成。我们设计了一个学习算法,保证了$\tilde{O}(\sqrt{T})$的regret界,这在时间范围$T$上至多有多项对数因子的紧致性。这与随机i.i.d.情况的minimax速率一致,并且与对抗性设置相分离,后者中sublinear-regret不可达。通过将i.i.d.情况下的强regret保证扩展到平滑对手,我们显著扩展了能够实现快速收敛的设置范围,同时填补了这个基本经济问题regret景观中的一个重要空白。为克服平滑对手带来的挑战,我们利用平滑实例的连续性性质,并结合经纪人的动作空间的分层网构造,通过算法链式分析进行分析。我们通过推导一个类似的紧致$\tilde{O}(\sqrt{T})$ regret界来展示这些技术的适用性,对于一个相关机制设计模型:联合广告问题。
英文摘要:
Bilateral trade models the task of intermediating between two strategic agents, a seller and a buyer, who wish to trade a good. We study this problem from the perspective of a profit-maximizing broker within an online learning framework, where the agents' valuations are generated by a smooth adversary. We devise a learning algorithm that guarantees a $\tilde{O}(\sqrt{T})$ regret bound, which is tight in the time horizon $T$ up to poly-logarithmic factors. This matches the minimax rate for the stochastic i.i.d. case, and is also well separated from the adversarial setting, where sublinear-regret is unattainable. By extending the strong regret guarantees from the i.i.d. case to the smooth adversary, we significantly broaden the scope of settings where such fast rate is achievable, while closing an important gap in the regret landscape of this fundamental economic problem. To overcome the challenges posed by this adversary, we leverage a continuity property of smooth instances and combines this with a hierarchical net-construction of the broker's action space, which is analyzed via algorithmic chaining. We showcase the applicability of these techniques by deriving a similarly tight $\tilde{O}(\sqrt{T})$ regret bound for a related mechanism design model: the joint ads problem.