面向贝叶斯非平稳老虎机的未来信息导向采样
Future Information-Directed Sampling for Bayesian Nonstationary Bandits
浏览论文内容
中文总结 AI 辅助
本文提出未来信息导向采样(FIDS),一种面向贝叶斯非平稳老虎机的新算法,通过显式探索未来最优臂信息,在保持与汤普森采样相近遗憾的同时利用预测性结构,并采用监督学习框架从离线数据近似后验推断。
中文摘要 AI 辅助
探索-利用是老虎机学习中的核心权衡。虽然经典算法如上置信界方法和汤普森采样在平稳环境中能有效平衡这一权衡,但其探索策略主要减少关于当前最优臂的不确定性,这在非平稳环境中可能不足,因为未来最优臂可能与当前最优臂存在显著差异。本文提出未来信息导向采样(FIDS),一种用于贝叶斯非平稳老虎机的新算法,该算法显式探索以收集关于未来最优臂的信息。我们证明,FIDS 的遗憾与汤普森采样相当,仅相差一个小常数因子,同时能够利用传统探索目标无法捕获的预测性信息结构。为解决后验推断的实际困难,我们进一步提出一种基于监督学习的近似框架,从离线数据中学习 FIDS 策略,并在合成基准上证明其有效性。
英文摘要
Exploration--exploitation is a central trade-off in bandit learning. While classical algorithms such as upper confidence bound methods and Thompson Sampling effectively balance this trade-off in stationary environments, their exploration strategies mainly reduce uncertainty about the current optimal arm, which can be insufficient in nonstationary settings where future optimal arms may differ substantially from current ones. In this paper, we propose Future Information-Directed Sampling (FIDS), a new algorithm for Bayesian nonstationary bandits that explicitly explores to gather information about future optimal arms. We show that FIDS achieves regret comparable to Thompson Sampling up to a small constant factor, while being able to exploit predictive information structures that conventional exploration objectives fail to capture. To address the practical difficulty of posterior inference, we further propose a supervised-learning-based approximation framework that learns the FIDS policy from offline data, and demonstrate its effectiveness on synthetic benchmarks.
发表机构
- Boston University(波士顿大学)
- Broad Institute of MIT and Harvard(麻省理工学院和哈佛大学布罗德研究所)
机构由 AI 辅助整理,请以论文原文为准。