专家问题与在线凸优化的极大极小交替遗憾
Minimax Alternating Regret for the Experts Problem and Online Convex Optimization
浏览论文内容
中文总结 AI 辅助
本文解决了专家问题与一般在线凸优化(OCO)中极大极小交替遗憾率的开放问题,给出了匹配的上下界,提出修正Hedge算法等技术,显著优于现有结果。
中文摘要 AI 辅助
受双人博弈中交替学习动态的成功启发,本文研究在线凸优化(OCO)中的交替遗憾。尽管已有研究表明,在损失函数和可行域的各种假设下可实现$o(\sqrt{T})$的交替遗憾,但即使对于专家问题,其极大极小遗憾率仍悬而未决。本文通过为专家问题和一般OCO匹配上下界解决了该问题。令人惊讶的是,对于$d$专家问题,极大极小交替遗憾为$\Theta(\log d)$,与时间范围$T$无关,这显著优于Hait等人[2025]确立的最佳已知$\mathcal{O}(T^{1/3}\log^{2/3} d)$。我们进一步将结果扩展到$d$维紧凸集上的一般OCO,证明最坏情况极大极小交替遗憾为$\Theta\left(d\log \left(1+\frac{T}{d}\right)\right)$,同样显著优于最佳已知$\mathcal{O}((d\log T)^{2/3}T^{1/3})$上界,并解决了Cevher等人[2023]、Hait等人[2025]提出的开放问题。技术上,我们为专家问题设计的上界通过Hedge的修正变体实现,其中精心设计的修正项抵消了交替遗憾分析中产生的不利曲率;我们将相同的修正势能论证扩展到连续动作集,以获得OCO的最优交替遗憾率。对于下界,专家构造反复剔除一半候选专家,而OCO下界实例构造则用单位圆盘上更复杂的多尺度构造替代这种离散剔除。
英文摘要
In this paper, we study alternating regret in online convex optimization (OCO), motivated by the success of alternating learning dynamics in two-player games. Although previous works have shown that $o(\sqrt{T})$ alternating regret is achievable under various assumptions on the loss functions and feasible domains, the minimax regret rate has remained open even for the expert problem. In this paper, we resolve this question by showing matching lower and upper bounds for both the expert problem and general OCO. Somewhat surprisingly, for the $d$-expert problem, we show that the minimax alternating regret is $Θ(\log d)$, independent of the horizon $T$. This significantly improves upon the best-known $\mathcal{O}(T^{1/3}\log^{2/3} d)$ established by Hait et al. [2025]. We further extend our results to general OCO over a $d$-dimensional compact convex set and prove that the worst-case minimax alternating regret is $Θ\left(d\log \left(1+\frac{T}{d}\right)\right)$, also significantly improving upon the best-known $\mathcal{O}((d\log T)^{2/3}T^{1/3})$ upper bound and resolving the open problem posed by Cevher et al. [2023], Hait et al. [2025]. Technically, our upper bound for the expert problem is achieved by a corrected variant of Hedge, in which carefully designed correction terms cancel the unfavorable curvature arising in the alternating-regret analysis. We extend the same corrected-potential argument to continuous action sets to obtain the optimal alternating-regret rate for OCO. For the lower bounds, the expert construction repeatedly eliminates half of the candidate experts, while the OCO lower bound instance construction replaces this discrete elimination by a more involved multiscale construction on the unit disk.
发表机构
- University of Iowa(爱荷华大学)
机构由 AI 辅助整理,请以论文原文为准。