延迟多臂老虎机中的池化与漂移
Pooling and Drift in Delayed Bandits
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对延迟多臂老虎机问题,利用行动产生的状态池化结果,提出旋转算法与单副本算法,证明其 regret 上界,经实验可显著降低 regret。
AI中文摘要:
系统往往需要在得知行动是否有效之前很久就采取行动:推荐系统在几秒内看到点击,而购买则在几天后发生。在有K个行动且延迟为d轮的情况下,已知该设置在T轮内的最佳速率为\\(\widetilde{O}(\sqrt{(K+d)T})\\),因此更长的行动菜单学习成本总是更高。但情况并非必然如此:如果结果仅通过行动产生的状态依赖于行动,那么一个延迟的结果会告知所有可能产生观察到的状态的行动,而成本由行动产生的真正不同状态的数量决定,而非行动的数量。我们使用介于1和状态数量之间的有效维度\\(v_t\\)来衡量这一点,并证明对于任意预先固定的预算,旋转算法的 regret 为\\(\widetilde{O}(\sqrt{(d+1)V\log K})\\),实际使用的单副本算法的 regret 为\\(\widetilde{O}(\sqrt{V^-}+\sqrt{dT})\\);合并相似状态会以明确的偏差进一步降低成本。即使获得d轮前的精确损失,也没有算法能突破\\(\Omega(\sqrt{dE\min\{1+\log J,T/d\}})\\),其中J是漂移方向的数量,E限定学习器等待时损失的移动幅度。在生成数据上,状态通道相比行动级加权将 regret 降低多达79%,在漏斗族上相比调优后的极小极大最优方法将 regret 降低32%至68%。
英文摘要:
A system often has to act long before it learns whether the act worked: a recommender sees a click in seconds and a purchase in days. With $K$ actions and a delay of $d$ rounds, the best rate known for this setting is $\widetilde{O}(\sqrt{(K+d)T})$ over $T$ rounds, so a longer menu is always more expensive to learn from. It need not be: if the outcome depends on the action only through the state it produced, then one late outcome informs every action that could have produced the observed state, and the price is set by how many genuinely different states the actions produce rather than by how many actions there are. We measure this using an effective dimension $v_t$ between $1$ and the number of states, and prove $\widetilde{O}(\sqrt{(d+1)V\log K})$ for a rotating algorithm and $\widetilde{O}(\sqrt{V^{-}}+\sqrt{dT})$ for the single-copy algorithm used in practice, for any budget fixed in advance; merging similar states lowers the price further, at an explicit bias. Even when given the exact losses from $d$ rounds ago, no algorithm escapes $Ω(\sqrt{dE\min\{1+\log J,T/d\}})$, where $J$ counts the drifting directions and $E$ bounds how far losses move while the learner waits. On generated data, the state channel cuts regret by up to 79 percent against action-level weighting and, on the funnel family, by 32 to 68 percent against a tuned minimax-optimal method.