arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

博弈中常数遗憾的乘性乐观方法

Multiplicative Optimism for Constant Regret in Games

Ashkan Soleymani, Georgios Piliouras

arXiv 2609.21976首次发表:更新:

发表机构

MIT; Google Deepmind(麻省理工学院; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出乘性乐观遗憾匹配(MORM),一种非耦合学习规则,在有限一般和博弈中实现常数遗憾,并扩展到对抗性效用场景。

AI 中文摘要

我们提出了乘性乐观遗憾匹配(MORM),一种用于有限一般和博弈的非耦合学习规则。在同时全信息自博弈下,每个玩家仅使用一步乐观即可在所有时间范围内均匀地实现外部遗憾 $O(\sqrt n\log d)$。该分析将基于势的遗憾匹配论证与乘性稳定性及策略移动的Hellinger控制相结合。此外,学习率保护机制在面对对抗性效用时提供了 $O(\sqrt{T\log d})$ 的遗憾。

英文摘要

We introduce Multiplicatively Optimistic Regret Matching (MORM), an uncoupled learning rule for finite general-sum games. Under simultaneous full-information self-play, every player achieves external regret $O(\sqrt n\log d)$ uniformly over all horizons, using only one-step optimism. The analysis combines a potential-based regret-matching argument with multiplicative stability and Hellinger control of strategy movement. A learning-rate safeguard additionally gives $O(\sqrt{T\log d})$ regret in the face of adversarial utilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑