在线上下文矩阵博弈的策略优化与统计推断
Policy Optimization and Statistical Inference for Online Contextual Matrix Games
浏览论文内容
中文总结 AI 辅助
针对策略收益随情境变化的在线博弈问题,提出在线上下文矩阵博弈框架与OnGameLearn算法,经模拟和酒店定价应用验证,可有效平衡探索利用并具备统计保证。
中文摘要 AI 辅助
在线决策通常需要在动态情境与策略互动共同塑造的环境中进行,例如在竞争性定价场景中,酒店必须同时考虑动态情境因素与竞争对手的策略反应。现有方法仅能应对部分挑战:上下文多臂老虎机利用可观测特征优化单智能体决策,但忽略多智能体互动;在线矩阵博弈通过纳什均衡捕捉策略行为,但假设收益固定,忽略情境信息。那么当策略收益随情境信号变化时,智能体应如何行动?我们提出在线上下文矩阵博弈,将情境信息整合至多智能体在线博弈中;进一步提出在线学习算法OnGameLearn,可在智能体行动与情境间高效平衡探索与利用。该方法具备统计保证:收益矩阵估计的尾界、估计纳什均衡的收敛性、参数估计量的渐近正态性以及次线性遗憾界。我们还定义了矩阵博弈中的策略价值概念,并提出其双稳健、√T一致的估计量。通过模拟研究与真实酒店定价应用,发现OnGameLearn可有效应对策略与情境决策交织的挑战。
英文摘要
Online decision making often requires navigating a landscape shaped by both dynamic contexts and strategic interactions. In competitive pricing, for example, hotels must account for both dynamic contextual factors and rivals' strategic responses. Existing approaches address only part of this challenge: contextual bandits optimize single-agent decisions using observable features but ignore multi-player interactions, while online matrix games capture strategic behavior through Nash equilibrium but assume fixed payoffs, ignoring contextual information. How should agents act then when strategic payoffs evolve with contextual signals? We introduce \emph{online contextual matrix games} to integrate contextual information into multi-player online games. We further propose \emph{OnGameLearn}, an online learning algorithm that efficiently balances exploration and exploitation across both player actions and contexts. This approach comes with statistical guarantees: tail bounds for the estimated payoff matrix, the convergence of the estimated Nash equilibrium, the asymptotic normality of the parameter estimators, and the sublinear regret bound. We also develop the notion of \emph{policy value} in matrix games and develop a doubly robust, $\sqrt{T}$-consistent estimator for it. Across simulated studies and a real-world hotel pricing application, we find that OnGameLearn effectively navigates the intertwined challenges of strategic and contextual decision-making.
发表机构
- University of California, Irvine(加州大学欧文分校)
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。