arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13886eess.SYcs.SY

部分可观测游荡多臂老虎机的最优阈值型策略

Optimal Threshold Type Policies for Partially Observable Restless Bandits

  • Department of Electrical Engineering, Indian Institute of Technology Madras(印度马德拉斯理工学院电气工程系)
  • Indian Institute of Information Technology, Allahabad(印度安拉阿巴德信息科技学院)

机构由 AI 辅助整理,请以论文原文为准。

Anu Krishna, Rahul Meshram, Kesav Ram Kaza

AI总结:

针对资源受限野生动物监测中的部分可观测游荡多臂老虎机问题,通过Lipschitz常数界定值函数变差,证明最优策略在信念单纯形上具有阈值型结构,并特化至一步生灭动态。

AI中文摘要:

我们研究了一个受资源受限的野生动物监测启发的有限状态部分可观测游荡多臂老虎机(PO-RMAB)问题。每个位置的基础状态独立演化,而在每个决策时刻只能主动监测有限数量的位置。主动激活会揭示当前状态,而被动操作则不提供任何观测,从而产生一种塌缩老虎机的信念动态。我们的主要贡献是对多维信念状态问题的最优策略进行结构性刻画。我们建立了充分条件,在这些条件下,激活优势相对于信念状态是单调的,因此最优策略在$(M-1)$维信念单纯形上具有阈值型结构。关键结果通过一个依赖于模型的Lipschitz常数来界定值函数的变差而获得。我们进一步将结果特化到一步生灭动态,并推导出显式界。

英文摘要:

We study a finite-state partially observable restless multi-armed bandit (PO-RMAB) motivated by resource-constrained wildlife monitoring. The underlying condition of each location evolves independently, while only a limited number of locations can be actively monitored at each decision epoch. Activation reveals the current state, whereas passive operation provides no observation, yielding a collapsing-bandit belief dynamics. Our main contribution is a structural characterisation of optimal policies for the multidimensional belief-state problem. We establish sufficient conditions under which an activation advantage is monotone with respect to the belief state, and hence the optimal policy has the threshold type structure over the $(M-1)$-dimensional belief simplex. The key result is obtained by bounding the variation of the value function through a model-dependent Lipschitz constant. We further specialise the result to one-step birth-death dynamics and derive explicit bounds.

↑