专家建议预测:多专家任意时刻遗憾匹配固定时间常数
Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant
浏览论文内容
中文总结 AI 辅助
本文针对专家建议预测问题,提出一种无需预知时间范围的算法,其任意时刻遗憾界匹配已知固定时间常数,消除了此前 $\sqrt{2}$ 的差距。
中文摘要 AI 辅助
专家建议预测是在线学习中的一个基本问题。当时间范围 $T$ 预先已知时,$n$ 个专家上的极小极大累积遗憾渐近为 $\sqrt{\frac{T \ln n}{2}}$。这一结果通过使用针对 $T$ 调整学习率的多重权重更新算法实现,并且已知是紧的。相反,如果遗憾界要求在每个时刻 $t$ 同时成立,则已知的最佳保证是 $\sqrt{t \ln n}$——比最优差 $\sqrt{2}$ 倍——并且尚不清楚这个 $\sqrt{2}$ 因子是否是必要的。我们证明它不是必要的。我们给出一种算法,无需知道时间范围,其累积遗憾满足 $R_t \le \bigl(1 + O(\sqrt{\ln \ln n / \ln n})\bigr)\sqrt{t \ln n / 2}$ 对每个 $t \ge 1$ 同时成立。
英文摘要
Prediction with expert advice is a fundamental problem in online learning. When the time horizon $T$ is known in advance, the minimax cumulative regret over $n$ experts is asymptotically $\sqrt{\frac{T \ln n}{2}}$. This is achieved by the Multiplicative Weights Update algorithm with a learning rate tuned to $T$, and is known to be tight. If instead the regret bound is required to hold simultaneously at every time $t$, the best known guarantee has been $\sqrt{t \ln n}$---a factor of $\sqrt{2}$ worse---and it has remained unknown whether this factor of $\sqrt{2}$ is necessary. We show that it is not. We give an algorithm, requiring no knowledge of the horizon, whose cumulative regret satisfies $R_t \le \bigl(1 + O(\sqrt{\ln \ln n / \ln n})\bigr)\sqrt{t \ln n / 2}$ simultaneously for every $t \ge 1$.
发表机构
- Google Research(谷歌研究院)
- Yale University(耶鲁大学)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。