发表机构
Google Research; Yale University; New York University; Google DeepMind(谷歌研究院; 耶鲁大学; 纽约大学; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出首个高效且适当的在线逆线性优化算法,实现 $O(d)$ 遗憾界和每轮 $O(d^2)$ 时间,通过自归一化更新与迹幂势函数改进,并具鲁棒性。
AI 中文摘要
我们提出了一种用于在线逆线性优化的确定性算法,其遗憾界为 $O(d)$,该界与时间范围无关,且每轮时间复杂度为 $O(d^{2})$。近期,Dewasurendra 获得了这一量级的界,解决了 Gollapudi 等人以及 Oki 和 Sakaue 提出的问题,但该算法是不适当的(improper),需要在每个尺度上枚举覆盖,且每轮代价为 $T^{\Theta(d)}$;我们的算法是首个高效且适当(proper)的此类界。我们基于 Sakaue 等人的变度量框架,加入了自归一化的秩一更新,并将 $\log\det$ 势函数替换为迹幂 $\tr(H^{-1/2})$,该量有界且消除了 $\ln T$ 项。该界同样适用于不进行优化的专家,我们还给出了抗腐败和秩自适应的变体,以及一个在凸最小化中的应用。
英文摘要
We give a deterministic algorithm for online inverse linear optimization with regret $O(d)$, uniform in the horizon and $O(d^{2})$ time per round. A bound of this order was obtained recently by Dewasurendra, settling a question of Gollapudi et al.\ and of Oki and Sakaue, but by an improper rule that enumerates covers at every scale and costs $T^{Θ(d)}$ a round; ours is the first efficient such bound and the first proper one. We build on the variable-metric framework of Sakaue et al., adding a self-normalized rank-one update, and we replace the $\log\det$ potential by the trace power $\tr(H^{-1/2})$, which is bounded outright and removes the $\ln T$. The bound also holds against an expert that does not optimize, and we give corruption-robust and rank-adaptive variants, and an application to convex minimization.