马尔可夫决策过程中的决策聚焦学习:一种占用度量方法
Decision-Focused Learning in MDPs: An Occupancy Measure Approach
- Johns Hopkins Carey Business School(约翰霍普金斯大学凯瑞商学院)
- International Computer Science Institute(国际计算机科学研究所)
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出基于占用度量线性规划的决策聚焦学习方法,通过枢轴算法求闭式梯度,结合增广拉格朗日与软状态聚合解决可扩展性问题,在多个任务上以更低计算成本实现更低遗憾。
AI中文摘要:
在本工作中,我们考虑针对马尔可夫决策过程(MDP)的决策聚焦学习(DFL),现有方法通过对贝尔曼方程的KKT条件进行微分,并需要求解覆盖所有状态-动作对的线性系统,这限制了其可扩展性。我们通过将MDP重新表述为基于占用度量的线性规划(LP)来解决这一问题,其可行域由预测的动态诱导,我们通过枢轴算法识别可行多面体中的活跃约束,从而推导出闭式梯度。这种基于占用度量的LP层带来了两个挑战:(1)当活跃约束变化时,LP的解梯度不连续;(2)LP的反向传播成本仍随状态规模增长,这对于大规模或连续状态空间而言代价高昂。我们通过增广拉格朗日代理来解决这些挑战,并通过随机行草图化约束来平滑边界跳跃,同时引入可学习的软状态聚合层及其函数逼近泛化,将LP扩展到大规模有限状态和连续状态的MDP。在多个任务中,我们的方法相比基于KKT的DFL和两阶段基线达到了更低的遗憾值,且计算成本显著降低。所有实验的源代码可在以下网址获取:https://this URL。
英文摘要:
In this work, we consider decision-focused learning (DFL) for a Markov decision process (MDP), where existing methods differentiate through the KKT conditions of the Bellman equation and require solving a linear system over all state-action pairs, limiting its scalability. We address this by reformulating the MDP as an occupancy measure-based linear program (LP), whose feasible region is induced by predicted dynamics, and we derive a closed-form gradient by identifying the active constraints in the feasible polyhedron via the pivoting algorithm. This occupancy measure-based LP layer raises two challenges: (1) LP's solution gradient is discontinuous when active constraints change, and (2) the LP backward cost still scales with the state size, which is costly for large or continuous state spaces. We address the challenges with an augmented Lagrangian surrogate and smooth the boundary jumps by random row sketching of the constraints, and a learnable soft state-aggregation layer and its function-approximation generalization that scales the LP to large finite and continuous-state MDPs. Across multiple tasks, our methods reach lower regret than KKT-based DFL and two-stage baselines with significantly lower computation cost. The source code for all experiments is available at https://github.com/A-Eshragh/State_Aggregation_Project.