arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.24234cs.LG

稀疏乐观信息导向采样

Sparse Optimistic Information Directed Sampling

  • Universitat Pompeu Fabra(庞培法布拉大学)

机构由 AI 辅助整理,请以论文原文为准。

Ludovic Schwartz, Hamish Flynn, Gergely Neu

更新

AI总结:

针对随机稀疏线性老虎机的高维在线决策问题,提出无需贝叶斯假设的SOIDS算法,通过时变学习率的新分析实现信息与遗憾的最优平衡,是首个在两种数据场景下均达最优最坏情况遗憾的算法,且实验性能优异。

AI中文摘要:

许多高维在线决策问题可建模为随机稀疏线性老虎机问题。现有大多数算法的设计目标是在两种场景下实现最优最坏情况遗憾:一是数据充足场景,此时对环境维度的多项式依赖不可避免;二是数据匮乏场景,此时可以实现维度无关,但代价是对轮数的依赖更差。相比之下,稀疏信息导向采样(IDS)算法满足的贝叶斯遗憾界能同时在两种场景下达到最优速率。本研究探索使用稀疏乐观信息导向采样(SOIDS),在无需贝叶斯假设的最坏情况设定下实现同样的自适应性。通过一种支持使用时变学习率的新颖分析,我们证明SOIDS能够最优地平衡信息与遗憾。我们的结果拓展了IDS的理论保证,提出了首个能同时在数据充足和数据匮乏场景下实现最优最坏情况遗憾的算法。我们通过实验验证了SOIDS的优异性能。

英文摘要:

Many high-dimensional online decision-making problems can be modeled as stochastic sparse linear bandits. Most existing algorithms are designed to achieve optimal worst-case regret in either the data-rich regime, where polynomial dependence on the ambient dimension is unavoidable, or the data-poor regime, where dimension-independence is possible at the cost of worse dependence on the number of rounds. In contrast, the sparse Information Directed Sampling (IDS) algorithm satisfies a Bayesian regret bound that has the optimal rate in both regimes simultaneously. In this work, we explore the use of Sparse Optimistic Information Directed Sampling (SOIDS) to achieve the same adaptivity in the worst-case setting, without Bayesian assumptions. Through a novel analysis that enables the use of a time-dependent learning rate, we show that SOIDS can optimally balance information and regret. Our results extend the theoretical guarantees of IDS, providing the first algorithm that simultaneously achieves optimal worst-case regret in both the data-rich and data-poor regimes. We empirically demonstrate the good performance of SOIDS.

↑