赌博机何时可以离开其锚点?非平稳性下经E-过程授权的汤普森采样
When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity
- UC Santa Barbara(加州大学圣塔芭芭拉分校)
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究非平稳多臂赌博机中何时允许遗忘,提出E-过程授权的汤普森采样(e-ATS),通过随时有效的E-过程控制折扣状态启用,在平稳性下以$\alpha_E$概率保持乐观采样,实验显示授权影响因环境而异。
AI中文摘要:
平稳性奖励记忆,但变化发生后,同样的历史可能会误导。我们探讨何时应允许遗忘。E-过程授权的汤普森采样(e-ATS)为每个臂提供全历史状态和折扣Beta状态。一个随时有效的E-过程首先授权折扣状态,然后一个可逆的相关性分数控制其影响。在授权之前,e-ATS完全遵循乐观汤普森采样(OTS)。在Beta-Bernoulli先验预测平稳模型下,e-ATS偏离OTS的概率至多为所选的$\alpha_E$,无需拟合阈值。相对于e-ATS,在注册套件上移除授权使平均归一化动态伪遗憾增加了$38.4\\%$,但在文献衍生的回放套件上却减少了$7.5\\%$。因此,证据控制适应何时开始,而非其是否总是有益。
英文摘要:
Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta-Bernoulli prior-predictive stationary model, e-ATS's probability of ever departing from OTS is at most the chosen $α_E$, without fitted thresholds. Relative to e-ATS, removing authorization increased mean normalized dynamic pseudo-regret by $38.4\%$ on the registered suite but reduced it by $7.5\%$ on the literature-derived replay suite. Therefore, evidence controls when adaptation begins, not whether it always helps.