协方差适应算法用于半带隙问题及其在稀疏奖励中的应用
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
- Inria Lille & ENS Paris-Saclay(法国里尔研究所与巴黎-萨克雷高等科学研究院)
- ENSAE & Criteo AI Lab(高等统计研究院与Criteo人工智能实验室)
- DeepMind Paris(伦敦DeepMind巴黎分公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究了随机组合半带隙问题,探讨了新的亚指数分布族,推导了更紧的期望遗憾下界,并构造了基于协方差估计的算法,应用于稀疏奖励场景。
AI中文摘要:
我们研究了随机组合半带隙问题,其中整个联合分布影响问题实例的复杂性(不同于标准带隙)。典型分布依赖于特定参数值,其先验知识在理论中是必需的,但在实践中难以估计;例如,通常假设的亚高斯族。我们通过考虑新的亚指数分布族来缓解这一问题,该族包含有界和高斯分布。我们证明了在该族上的新期望遗憾下界,该下界参数化于结果的未知协方差矩阵,比亚高斯矩阵更紧。然后构造了一个使用协方差估计的算法,并提供了一个紧致的渐近分析。最后,我们将结果扩展到稀疏结果族,这在许多推荐系统中有应用。
英文摘要:
We investigate stochastic combinatorial semi-bandits, where the entire joint distribution of outcomes impacts the complexity of the problem instance (unlike in the standard bandits). Typical distributions considered depend on specific parameter values, whose prior knowledge is required in theory but quite difficult to estimate in practice; an example is the commonly assumed sub-Gaussian family. We alleviate this issue by instead considering a new general family of sub-exponential distributions, which contains bounded and Gaussian ones. We prove a new lower bound on the expected regret on this family, that is parameterized by the unknown covariance matrix of outcomes, a tighter quantity than the sub-Gaussian matrix. We then construct an algorithm that uses covariance estimates, and provide a tight asymptotic analysis of the regret. Finally, we apply and extend our results to the family of sparse outcomes, which has applications in many recommender systems.