arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一般博弈中的恒定个体遗憾

Constant Individual Regret in General Games

Mingyang Liu, Gabriele Farina, Asuman Ozdaglar

arXiv 2608.31166首次发表:更新:

发表机构

Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对完全信息反馈下的有限N玩家标准型博弈,提出确定性完全非耦合的ECHO-OFTRL算法,消除个体遗憾与时间跨度的多对数依赖,保证各玩家遗憾上界为O(poly(N, log m_max))。

AI 中文摘要

非耦合无遗憾动态机制为均衡提供了去中心化路径,但先前关于个体遗憾的保证仍与时间跨度呈多对数依赖关系。针对完全信息反馈下的每个有限N玩家标准型博弈,我们消除了这种依赖关系。我们引入ECHO-OFTRL:配备用于高阶乐观的EMA级联(ECHO,EMA指指数移动平均)的乐观正则化跟随领导者(OFTRL)算法,该算法为确定性且完全非耦合。若m_max表示最大动作集大小,则对于每个时间跨度T≥1,它保证博弈中N个玩家各自的遗憾上界为O(poly(N, log m_max))。我们的算法利用了受现代滤波器设计启发的新型乐观形式。

英文摘要

Uncoupled no-regret dynamics provide a decentralized route to equilibrium, but prior guarantees for individual regret retain a polylogarithmic dependence on the horizon. We remove this dependence for every finite $N$-player normal-form game under full-information feedback. We introduce \emph{ECHO-OFTRL}: optimistic follow-the-regularized-leader (OFTRL) equipped with an EMA cascade for high-order optimism (ECHO), where EMA denotes exponential moving average. The algorithm is deterministic and fully uncoupled. If $m_{\max}$ denotes the largest action-set size, then, simultaneously for every horizon $T\geq1$, it guarantees that each of the $N$ players in the game incurs regret upper bounded by $O(\textrm{poly}(N, \log m_{\max}))$. Our algorithm leverages a new form of optimism inspired by modern filter design.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑