arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩展式博弈中的最优策略追踪

Tracking the Best Strategy in an Extensive-Form Game

Stephen Pasteris, Rahul Savani, Theodore Turocy

arXiv 2608.09501首次发表:更新:

发表机构

The Alan Turing Institute; The University of Liverpool; The University of East Anglia(阿兰·图灵研究所; 利物浦大学; 东英吉利大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩展式博弈场景,提出了一种高效算法,可实现特定复杂度的切换悔惜,用于追踪最优策略。

AI 中文摘要

我们研究扩展式多臂老虎机问题,其中学习者在每一轮对局中都会与一个非适应性对手进行一场扩展式博弈。我们聚焦于切换悔惜的概念,该概念用于衡量学习者的期望表现与任意混合策略切换序列的事后表现之间的差距。我们的算法以ρ>0为参数,实现了切换悔惜的复杂度为$\tilde{\u2134}((1/\rho+\rho K)\sqrt{H A T})$,其中K为对比序列中的切换次数,H为对局中学习者可遍历的最大信息集数量,A为学习者可能采取的动作数量。该算法效率极高,每轮对局的时间复杂度仅为$\u2134(H B)$,其中B为学习者在任意信息集下可采取的最大动作数量。

英文摘要

We consider the extensive-form bandit problem where on each trial the learner plays an extensive-form game against an oblivious adversary. We focus on the notion of switching regret, which measures the expected performance of the learner against that of any switching sequence of mixed strategies in retrospect. Our algorithm takes a parameter $ρ>0$ and achieves a switching regret of $\tilde{\mathcal{O}}((1/ρ+ρK)\sqrt{H A T})$ where $K$ is the number of switches in the comparator sequence, $H$ is the maximum number of the learner's information sets that can be traversed during a play of the game and $A$ is the number of actions that the learner can possibly take. Our algorithm is extremely efficient, taking a per trial time of only $\mathcal{O}(H B)$ where $B$ is the maximum number of actions available to the learner at any of its information sets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑