AI 中文总结
该研究提出一种两人零和重复博弈,其值恒等式关联贝叶斯更新与指数权重遗憾,分解出三类遗憾项,且多类学习方法均为其特例。
AI 中文摘要
我们提出了一个学习者与自然之间的两人零和重复博弈,其值恒等式可同时生成贝叶斯更新和指数权重遗憾的精确计算,并提供了广泛浓度现象共享的比较器类变分形式。终端收益是比较器在相对于先验的固定相对熵下能获得的最大值,而一步约束是自然在学习者混合行动下的行动信息预算。在学习者行动无其他限制的情况下,Gibbs/Bayes权重会作为其唯一的Bellman均衡出现——即使得每轮损失与自然行动方向无关的混合行动,其中对数配分函数扮演值函数的角色。遗憾可精确分解为三部分:反映观测结果变化的每轮信息损失、精确解释轮次间测量尺度任何变化的附加重调漂移项,以及比较器相对于先验携带的信息。驱动标准遗憾界的方差和有界范围代理是该分解的更宽松松弛,该分解普遍成立并支配所有这些代理。双方的策略可逐项从分解中推导得出,重复博弈会产生自博弈的信息论账本,替代通常的二次变差代理。相同的比较器类几何解释了经典大偏差界,多臂老虎机、后验采样、聚合和提升方法均是该遗憾分解的特例。
英文摘要
We give a two-player zero-sum repeated game between a learner and nature whose value identity generates Bayesian updating and an exact accounting of exponential-weights regret at once, and supplies the comparator-class variational form that a wide class of concentration phenomena share. The terminal payoff is the most a comparator can gain at fixed relative entropy from the prior, and the one-step constraint is an information budget on nature's move under the learner's mixed action. With the learner's move otherwise unrestricted, Gibbs/Bayes weights emerge as its unique Bellman equalizer -- the mixed action that makes the per-round loss independent of which direction nature moves -- with log-partition functions playing the role of value functions. The regret decomposes exactly into three parts: a per-round information loss reflecting the variation in observed outcomes, an additive retempering drift that accounts exactly for any change of measurement scale between rounds, and the information the comparator carries relative to the prior. The variance and bounded-range proxies that drive standard regret bounds are looser relaxations of this decomposition, which holds generally and governs them all. Both players' strategies are read off from the decomposition term by term, and repeated play yields an information-theoretic ledger of self-play in place of the usual quadratic-variation surrogate. The same comparator-class geometry accounts for the classical large-deviation bounds, and methods across bandits, posterior sampling, aggregation, and boosting are specializations of the one regret decomposition.