AI 中文总结
针对近似线性规划权重选择的问题,提出基于对偶解投影占用信息的自适应权重更新方法,在更低计算成本下提升策略质量,且收敛性有保障。
AI 中文摘要
近似线性规划(ALP)广泛应用于大规模马尔可夫决策过程(MDP),但其性能对状态相关性权重的选择十分敏感,这类权重通常采用启发式方法选取。性能界表明,应使这些权重与诱导策略的折扣占用测度对齐,现有原始方法通过反复构造贪心策略实现这一点,但缺乏收敛性保证且计算成本高昂。本文提出一种基于对偶的方法,利用ALP对偶解的投影占用信息构造平滑随机策略并更新状态相关性权重,避免了单独的贪心动作计算。我们建立了权重与诱导策略折扣占用测度匹配的条件,并证明在适当平滑下的唯一性和全局收敛性。此外,我们推导了后验策略损失界,该损失界将误差从加权贝尔曼残差、占用不匹配以及随机与贪心策略的分歧中分离出来。在经典排队和多优先级调度问题上的实验表明,所提方法降低了对固定权重的敏感性,在计算成本更低的情况下,实现了与原始方法相当或更优的策略质量。最后,我们证明当基函数足够能使占用信息影响最终策略时,自适应权重最具价值。
英文摘要
Approximate Linear Programming (ALP) is widely used for large-scale Markov Decision Processes (MDPs), but its performance can be sensitive to the choice of state-relevance weights, which are typically selected heuristically. Performance bounds suggest aligning these weights with the discounted occupancy measure of the induced policy, and existing primal approaches address this through repeated greedy-policy construction. Nonetheless, they lack convergence guarantees and are computationally expensive. We propose a dual-based method that uses projected occupancy information from the ALP dual solution to construct a smooth stochastic policy and update the state-relevance weights, which avoids separate greedy-action calculations. We establish conditions under which the weights match the discounted occupancy of the induced policy and prove uniqueness and global convergence under appropriate smoothing. We also derive an a posteriori policy-loss bound that separates error from the weighted Bellman residual, occupancy mismatch, and stochastic-versus-greedy disagreement. Experiments on classical queueing and multi-priority scheduling problems show that the proposed approach reduces sensitivity to fixed weights and achieves comparable or better policy quality than primal updates at lower computational cost. Finally, we show that adaptive weighting is most valuable when the basis functions are sufficiently expressive for occupancy information to influence the resulting policy.
Comments20 pages, 9 figures