arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

有限时间范围老虎机问题中的贪心优势

The Greedy Advantage in Finite-Horizon Bandits

Kai Zhou, Michael Lingzhi Li, Kai Wang

arXiv 2607.29375首次发表:更新:

发表机构

Tsinghua University; Harvard Business School(清华大学; 哈佛商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对有限时间范围老虎机问题,提出正则化贪心算法,推导其有限时间范围遗憾包络并校准参数,实验显示其性能优于或匹配当前最优算法,为该问题提供有效方案。

AI 中文摘要

各机构越来越依赖序贯实验来优化决策。多臂老虎机文献已提出具有强渐近遗憾保证的算法,但许多实际应用在有限且外部给定的时间范围内运行。受有限时间范围设定的启发,我们针对多臂伯努利老虎机开发了一类正则化贪心算法。我们推导了正则化贪心老虎机的首个有限时间范围遗憾包络,表明有限时间范围遗憾可分解为暂态探索成本和次优收敛项,后者随正则化强度呈指数衰减。该特性为正则化参数提供了原则性校准规则,且作为极限情况,为经典贪心策略提供了更精确的遗憾保证。在大量数值实验中,经校准的正则化贪心策略始终与当前最优算法表现相当或更优。这些结果表明,正则化贪心策略可为有限时间范围老虎机问题提供有效解决方案。

英文摘要

Organizations increasingly rely on sequential experimentation to improve decision-making. While the multi-armed bandit literature has developed algorithms with strong asymptotic regret guarantees, many practical applications operate over finite and externally imposed horizons. Motivated by the finite-horizon setting, we develop a class of regularized greedy algorithms for multi-armed Bernoulli bandits. We derive the first finite-horizon regret envelopes for regularized greedy bandits, showing that finite-horizon regret decomposes into transient exploration costs and a suboptimal convergence term that decays exponentially with the regularization strength. This characterization yields principled calibration rules for the regularization parameters and, as a limiting case, sharper regret guarantees for the classical greedy policy. Across extensive numerical experiments, calibrated regularized greedy policies consistently match or outperform state-of-the-art algorithms. These results suggest that regularized greedy policies can provide an effective approach for finite-horizon bandit problems.

Comments122 pages, 3 figures, submitted to Management Science and under peer review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑