arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2312.07145cs.LGstat.ML

基于在线神经回归的上下文老虎机

Contextual Bandits with Online Neural Regression

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

Rohan Deb, Yikun Ban, Shiliang Zuo, Jingrui He, Arindam Banerjee

更新

AI总结:

本研究针对上下文老虎机问题,通过在宽网络预测中加入随机扰动使损失满足QG条件,得到了更优的NeuCB遗憾界,并通过实验验证了算法(尤其是KL损失版本)的性能优势。

AI中文摘要:

近期研究已证明在可实现性假设下,上下文老虎机可归约为在线回归问题[Foster and Rakhlin, 2020, Foster and Krishnamurthy, 2021]。本研究探究了将神经网络用于此类在线回归及相关的神经上下文老虎机(Neural Contextual Bandits, NeuCBs)的效果。利用宽网络的现有结论,可直接推导出平方损失下在线回归的遗憾界为${\tanh{O}}(\tanh{\tanh{T}})$,通过归约可进一步得到NeuCBs的遗憾界为${\tanh{O}}(\tanh{K} T^{3/4})$。\n\n不同于这一标准方法,我们首先证明:对于满足QG(二次增长,Quadratic Growth)条件——PL(Polyak-Łojasiewicz)条件的推广——且具有唯一最小值的几乎凸损失,在线回归的遗憾界为$\tanh{O}(\tanh T)$。尽管宽网络因不具备唯一最小值而无法直接应用该结论,但我们发现,向网络预测结果中添加合适的微小随机扰动后,损失会出人意料地满足带唯一最小值的QG条件。基于这种扰动预测,我们证明了平方损失和KL损失下在线回归的遗憾界均为${\tanh{O}}(\tanh T)$,并进一步将其分别转化为NeuCB的$\tilde{\tanh{O}}(\tanh{KT})$和$\tilde{\tanh{O}}(\tanh{KL^*} + K)$遗憾界,其中$L^*$为最优策略的损失。\n\n此外,我们还证明,现有NeuCB的遗憾界要么为$\tanh{\tanh}(T)$,要么假设上下文独立同分布,均与本研究的结论不同。最后,我们在多个数据集上的实验结果表明,所提出的算法——尤其是基于KL损失的算法——性能始终优于现有算法。

英文摘要:

Recent works have shown a reduction from contextual bandits to online regression under a realizability assumption [Foster and Rakhlin, 2020, Foster and Krishnamurthy, 2021]. In this work, we investigate the use of neural networks for such online regression and associated Neural Contextual Bandits (NeuCBs). Using existing results for wide networks, one can readily show a ${\mathcal{O}}(\sqrt{T})$ regret for online regression with square loss, which via the reduction implies a ${\mathcal{O}}(\sqrt{K} T^{3/4})$ regret for NeuCBs. Departing from this standard approach, we first show a $\mathcal{O}(\log T)$ regret for online regression with almost convex losses that satisfy QG (Quadratic Growth) condition, a generalization of the PL (Polyak-Łojasiewicz) condition, and that have a unique minima. Although not directly applicable to wide networks since they do not have unique minima, we show that adding a suitable small random perturbation to the network predictions surprisingly makes the loss satisfy QG with unique minima. Based on such a perturbed prediction, we show a ${\mathcal{O}}(\log T)$ regret for online regression with both squared loss and KL loss, and subsequently convert these respectively to $\tilde{\mathcal{O}}(\sqrt{KT})$ and $\tilde{\mathcal{O}}(\sqrt{KL^*} + K)$ regret for NeuCB, where $L^*$ is the loss of the best policy. Separately, we also show that existing regret bounds for NeuCBs are $Ω(T)$ or assume i.i.d. contexts, unlike this work. Finally, our experimental results on various datasets demonstrate that our algorithms, especially the one based on KL loss, persistently outperform existing algorithms.

↑