arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

隐藏曲率的代价:带反馈凸优化的\(\widetilde{\Omega}(d^{5/4}\sqrt{T})\)下界

The Price of Hidden Curvature: Improved Lower Bounds for Bandit Convex Optimization

Nived Rajaraman, Yanjun Han

arXiv 2607.18652首次发表:更新:

发表机构

Microsoft Research; New York University(微软研究院; 纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究随机带反馈凸优化中1-利普希茨函数的极小极大期望遗憾下界,构造特殊凸函数类,通过费希尔信息矩阵后验扩散分析,得出\(\widetilde{\Omega}(d^{5/4}\sqrt{T})\)下界并扩展到无约束设置,揭示其比线性带反馈优化更难。

AI 中文摘要

我们建立了一个关于欧几里得球上1-利普希茨函数的随机带反馈凸优化的极小极大期望遗憾的\(\widetilde{\Omega}(d^{5/4}\sqrt{T})\)下界。这是该问题第一个比\(d\sqrt{T}\)增长更快的非平凡遗憾下界,表明随机带反馈凸优化从根本上比线性带反馈优化更难。我们构造的难凸函数类在维度\(2d\)下具有特定形式。观测仅在学习者动作靠近由\(W^\star\)确定的管时才提供关于\(u^\star\)的信息。我们的遗憾分析通过限制在自适应动作序列下获得的费希尔信息矩阵的后验扩散来利用这种权衡。这给出了找到\(\varepsilon\)-最优动作的样本复杂度下界\(\widetilde{\Omega}(d^{5/2}/\varepsilon^2)\),转化为\(\widetilde{\Omega}(d^{5/4}\sqrt{T})\)遗憾下界。我们还将此下界扩展到动作空间为\(\mathbb{R}^d\)的无约束设置。

英文摘要

We establish improved lower bounds on the minimax expected regret of stochastic bandit convex optimization for $1$-Lipschitz functions on the $d$-dimensional Euclidean ball. For time horizons $n\ge d^{10/3}$, we prove a lower bound of $Ω(d^{4/3}\sqrt{n})$, the first nontrivial bound that exceeds the $d\sqrt{n}$ dependence of linear bandits, showing that stochastic bandit convex optimization is fundamentally harder than linear bandits. For $d^2\le n\le d^{10/3}$, we obtain a lower bound of $Ω(\sqrt{d}n^{3/4})$, matching the regret of the algorithm of Flaxman et al. (2005), establishing its optimality in this regime. The hard class of convex functions we construct takes the following form in dimension $2d$: for an action $a=(a^1,a^2)\in \mathbb{B}^{2d}$, each function is the scaled soft maximum of a "tube", $r^{-1}\|W^\star a^1-\frac{r}{8\varepsilon}a^2 \|$ (hyperparameterized by $\varepsilon,r$), and a squared distance function, $\frac12\|a^1-u^\star\|^2-\frac12\|u^\star\|^2$. Here $u^\star\in\mathbb{R}^d$ is the unknown target determining the minimizer, while $W^\star\in\mathbb{R}^{d\times d}$ hides the region in which the quadratic curvature is observable. Indeed, observations reveal substantial information about $u^\star$ only when the learner acts near the hidden tube $a^2\approx \frac{8\varepsilon}{r}W^\star a^1$; away from it, the tube branch masks the quadratic branch. Thus the learner must pay to uncover the geometry encoded by $W^\star$ before it can effectively exploit the curvature that identifies $u^\star$. Formalizing this tradeoff yields a sample complexity lower bound of $Ω(\frac{d^{5/2}}{\varepsilon^2}\wedge\frac{d^2}{\varepsilon^4})$ for finding an $\varepsilon$-optimal action, and ultimately the $Ω(d^{4/3}\sqrt{n}\wedge\sqrt{d}n^{3/4})$ regret lower bound. The proof was developed by GPT-5.5 Pro and GPT-5.6 Sol Pro under the authors' guidance.

Comments44 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑