An optimal algorithm for bandit convex optimization
- Princeton University(普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。
更新
英文摘要:
We consider the problem of online convex optimization against an arbitrary adversary with bandit feedback, known as bandit convex optimization. We give the first $\tilde{O}(\sqrt{T})$-regret algorithm for this setting based on a novel application of the ellipsoid method to online learning. This bound is known to be tight up to logarithmic factors. Our analysis introduces new tools in discrete convex geometry.