AI 中文总结
针对周期性非平稳多臂老虎机问题,提出周期性自展汤普森采样(PBTS),通过同步信念重置、嵌入自展探索阶段克服传统TS缺陷。实验表明PBTS能显著降低累积遗憾,还阐述了其实际部署潜力与局限及未来研究方向。
AI 中文摘要
本文介绍了周期性自展汤普森采样(PBTS),它是经典汤普森采样(TS)算法的创新扩展,专为具有周期性非平稳性的老虎机问题量身定制。传统TS累积所有过去观测值,在奖励分布随时间循环时会导致后验偏差。PBTS通过将信念重置与已知或推断的周期间隔同步,并嵌入结构化自展探索阶段来克服这一问题,有效清除过时数据同时保留不确定性估计。PBTS在人工构建的环境中进行测试,包括偏态和平衡奖励分布、不同自展比例和未对齐的周期间隔。结果表明,在周期性非平稳环境中,PBTS通常在累积遗憾方面比传统TS有统计学显著降低。后续讨论进一步阐明了PBTS实际部署的潜力。研究提到了极端周期未对齐等局限性,并提出了如自调整周期识别等未来研究方向。通过内存重置和自展阶段,PBTS为周期性奖励环境中的老虎机算法优化引入了一种新方法。
英文摘要
This paper introduces Periodic Bootstrap Thompson Sampling (PBTS), an innovative extension of the classic Thompson Sampling (TS) algorithm tailored for bandit problems with periodic non-stationarity. Conventional TS accumulates all past observations, leading to biased posteriors when reward distributions cycle over time. PBTS overcomes this by synchronizing belief resets with known or inferred period intervals and embedding structured bootstrap exploration phases, effectively purging obsolete data while preserving uncertainty estimates. PBTS is tested in artificially constructed environments, which include skewed and balanced reward distributions, along with different bootstrap proportions and misaligned periodic intervals. Results indicate that PBTS generally achieves statistically significant reductions in cumulative regret against traditional TS in periodic non-stationary environments. Subsequent discussion further articulates the potential of PBTS's real-world deployment. The study mentions limitations like extreme periodic misalignment and proposes future research such as self-adjusting cycle-recognition. With memory reset and bootstrap phase, PBTS introduces a novel approach to optimizing bandit algorithms in periodic reward contexts.
Comments9 pages, 10 figures, Accepted to 2025 3rd International Conference on Data Science, Advanced Algorithms, and Intelligent Computing