arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18863stat.MLcs.LG

采用恒定探索的时变高斯过程背包问题的更紧遗憾界

Sharper Regret Bounds for Time-Varying Gaussian Process Bandits with Constant Exploration

  • Delft Institute of Applied Mathematics(代尔夫特应用数学研究所)
  • Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Matthias Mandl, Hanne Kekkonen

AI总结:

针对时变高斯过程漂移模型下的贝叶斯优化问题,提出基于每轮局部置信事件的恒定探索GP-UCB方法,推导更紧的时变最大信息增益界与期望、实际遗憾界,仿真验证了探索参数与漂移率倒数的对数依赖关系。

AI中文摘要:

我们研究时变环境下的贝叶斯优化,该环境中未知奖励函数按照高斯过程漂移模型演化。现有该场景下的GP-UCB(高斯过程置信上界)分析通常要求探索参数随时间步长增长以维持一致置信界。利用每轮局部置信事件,我们证明GP-UCB可采用恒定探索参数运行,并得到系数由漂移率决定的期望遗憾界。我们还推导了更紧的时变最大信息增益界。针对平方指数核,在持续漂移情形下,该界给出$\tildeγ_T/T=\tilde{\fancyscript{O}}(ε^{1/2})$以及$\tilde{\fancyscript{O}}(ε^{1/4})$的期望平均遗憾。相同的恒定探索分析还给出了实际遗憾保证。仿真结果验证了界所建议的探索参数与$1/ε$的对数依赖关系。

英文摘要:

We study Bayesian optimization in a time-varying environment where the unknown reward function evolves according to a Gaussian process drift model. Existing GP-UCB analyses in this setting typically require the exploration parameter to grow with the horizon to maintain uniform confidence bounds. Using per-round local confidence events, we show that GP-UCB can instead be run with a constant exploration parameter and obtain an expected-regret bound whose coefficient depends on the drift rate. We also derive a sharper time-varying maximum-information-gain bound. For the squared exponential kernel, it yields $\tildeγ_T/T=\widetilde{\mathcal O}(ε^{1/2})$ and expected average regret $\widetilde{\mathcal O}(ε^{1/4})$ in the persistent-drift regime. The same constant-exploration analysis also yields realized-regret guarantees. Simulations support the predicted logarithmic dependence of the bound-suggested exploration parameter on $1/ε$.

补充信息

↑