AI 中文总结
研究贝叶斯和乘法权重更新的遗憾问题,提出基于信息核算恒等式的方法,给出累积遗憾的精确自适应分解,该演算涵盖多种学习和博弈场景,揭示有利情况的自界属性。
AI 中文摘要
贝叶斯和乘法权重更新根据顺序反馈对专家、模型或行动进行重新加权。我们表明,任何此类更新的遗憾都遵循精确的信息核算恒等式。在每一轮中,学习者相对于任何选定比较器的超额损失是该轮暴露的不确定性的即时支付与学习者当前权重到比较器的信息距离减少量之和。累积支付定义了一个逐路径的不确定性时钟,即实现序列的“内在时间”。对单步平衡求和产生了累积遗憾的两个精确自适应分解,每种分解对应一种跨轮组合更新的自然方式。由于这些分解是精确的而非上界,有利的随机或低噪声情况表现为实现的内在时间的自界属性,而非最坏情况分析中的松弛。相同的演算涵盖了对冲、乐观和附带信息变体、连续先验、增强学习、在线凸优化、上下文博弈和重复博弈:每种情况下的逐路径核算都是相同的。
英文摘要
Bayesian and multiplicative-weights updates reweight experts, models, or actions from sequential feedback. We show that the regret of any such update obeys an exact information-accounting identity. On each round, the learner's excess loss to any chosen comparator is the sum of an immediate cost for the uncertainty exposed by the round and a reduction in the information distance from the learner's current weights to the comparator. The cumulative cost defines a pathwise uncertainty clock, the intrinsic time of the realized sequence. Summing one-step balances yields two exact adaptive decompositions of cumulative regret, one for each natural way of composing the update across rounds. Because the decompositions are exact, favorable stochastic or low-noise regimes appear as self-bounding properties of the realized intrinsic time. The accounting also fixes a learning rate, inverse in the square root of intrinsic time. That schedule is competitive with adaptive baselines in selected online-learning settings. The same calculus covers Hedge, optimistic and side-information variants, continuous priors, boosting, online convex optimization, contextual bandits, and repeated games: the pathwise account is the same in every case.