arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.20096cs.LGcs.ITmath.IT

信息导向采样用于老虎机:入门

Information-directed sampling for bandits: a primer

  • Institute of Physics, School of Basic Sciences, École Polytechnique Fédérale de Lausanne - EPFL, Lausanne, Switzerland
  • Quantitative Life Sciences section, The Abdus Salam International Center for Theoretical Physics (ICTP), Trieste, Italy
  • Department of Oncology, Universit\`a degli Studi di Torino, Italy

机构由 AI 辅助整理,请以论文原文为准。

Annika Hirling, Giorgio Nicoletti, Antonio Celani

更新

AI总结:

本文介绍信息导向采样策略在老虎机问题中的应用,通过分析对称老虎机和一枚公平硬币场景,展示IDS在不同设定下的性能表现。

AI中文摘要:

多臂老虎机问题提供了一个分析探索与利用之间张力的基本框架。本文探讨了信息导向采样(IDS)策略,这一类启发式方法在即时遗憾与信息增益之间取得平衡。我们聚焦于两状态伯努利老虎机的可处理环境,作为最小模型来严格比较启发式策略与最优策略。我们通过引入修改的信息度量和调节参数,将IDS框架扩展到折扣无限 horizon 设置。我们考察了两个具体问题类别:对称老虎机和涉及一枚公平硬币的场景。在对称情况下,我们证明IDS实现了有界的累积遗憾,而在一枚公平硬币的情况下,IDS策略产生的遗憾随着 horizon 的增长呈对数比例增长,与经典渐近下界一致。本文旨在作为教学性综合,旨在为统计物理学家的受众架起强化学习和信息论的概念桥梁。

英文摘要:

The Multi-Armed Bandit problem provides a fundamental framework for analyzing the tension between exploration and exploitation in sequential learning. This paper explores Information Directed Sampling (IDS) policies, a class of heuristics that balance immediate regret against information gain. We focus on the tractable environment of two-state Bernoulli bandits as a minimal model to rigorously compare heuristic strategies against the optimal policy. We extend the IDS framework to the discounted infinite-horizon setting by introducing a modified information measure and a tuning parameter to modulate the decision-making behavior. We examine two specific problem classes: symmetric bandits and the scenario involving one fair coin. In the symmetric case we show that IDS achieves bounded cumulative regret, whereas in the one-fair-coin scenario the IDS policy yields a regret that scales logarithmically with the horizon, in agreement with classical asymptotic lower bounds. This work serves as a pedagogical synthesis, aiming to bridge concepts from reinforcement learning and information theory for an audience of statistical physicists.

↑