AI 中文总结
针对部分可观测不安分赌机,提出t步前瞻阈值策略,可验证可索引性,其近似惠特尔指数几何收敛,性能优于一步基线且接近最优基准。
AI 中文摘要
惠特尔指数策略为不安分多臂赌机提供了一种可扩展方法,但在部分可观测情况下,即使确定单个信念处的无差异补贴也需求解无限 horizon 信念状态问题,且无闭式价值函数。Liu[10]通过线性化未知决策边界解决该难题,得到线性方程组与闭式近似惠特尔指数,但所得阈值仅采用一步主动-被动比较,未考虑更长 horizon 的延续价值。我们将该框架扩展为t步前瞻阈值策略:对每个补贴m,阈值由t步有限 horizon 价值迭代下的主动减被动优势定义。当t=1时,阈值与m无关,恢复Liu[10]的线性阈值;当t>1时,阈值通过诱导的首交叉结构变为补贴依赖,更紧密跟踪精确决策边界。所提算法无需将可索引性作为输入,且包含可索引性验证。在原始惠特尔可索引性下,我们证明t步近似惠特尔指数几何收敛到精确惠特尔指数,即|Ŵ_t(ω)-W(ω)|=O(β^t)。数值实验中,按所提准则验证的2715个测试三状态实例均为可索引;P95指数误差从t=1时的2.18×10^-2降至t=8时的8.93×10^-4;在β=0.9999的精确可比实例中,t=2已恢复精确惠特尔指数排序;中等深度阈值策略优于一步基线,且接近最优动态规划基准,运行时间随t增长缓慢。
英文摘要
Whittle index policies offer a scalable method for restless multi-armed bandits, but under partial observability even determining the indifference subsidy at a single belief requires solving an infinite-horizon belief-state problem with no closed-form value function. Liu [10] addresses this difficulty by linearizing the unknown decision boundary, leading to a linear system and a closed-form approximate Whittle index. However, the resulting threshold uses only a one-step active--passive comparison and does not account for longer-horizon continuation values. We extend this framework to a \emph{$t$-step lookahead threshold policy}. For each subsidy $m$, the threshold is defined by the active-minus-passive advantage under $t$-step finite-horizon value iteration. At $t=1$, the threshold is $m$-independent and recovers the linear threshold of Liu [10]; for $t>1$, it becomes subsidy-dependent through the induced first-crossing structure and tracks the exact decision boundary more closely. The proposed algorithmic framework does not require indexability as an input and includes an indexability verification. Under the original Whittle indexability, we prove that the $t$-step approximate Whittle index converges geometrically to the exact Whittle index, \[ |\widehat W_t(ω)-W(ω)|=O(β^t). \] Numerically, all 2,715 tested three-state instances are verified with computable priority index functions according to the proposed criterion. The P95 index error decreases from $2.18\times10^{-2}$ at $t=1$ to $8.93\times10^{-4}$ at $t=8$. In an exact-comparable instance with $β=0.9999$, $t=2$ already recovers the exact Whittle-index ordering. Moderate-depth threshold policies also outperform the one-step baseline and remain close to the optimal dynamic-programming benchmark, while runtime grows mildly with $t$.
CommentsFixed some errors in the previous version