硬约束、平滑梯度:通过可微投影学习可行库存策略
Hard Constraints, Smooth Gradients: Learning Feasible Inventory Policies via Differentiable Projection
- TUM School of Management, Technical University of Munich(慕尼黑工业大学TUM管理学院)
- Esade, Ramon Llull University(拉蒙·柳利大学ESADE商学院)
- Munich Data Science Institute (MDSI), Technical University of Munich(慕尼黑工业大学慕尼黑数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出嵌入可微凸优化模块的DRL策略,以处理序列决策中的硬约束,在库存规划等场景中实现了优于基准方法的成本降低与最优性表现。
AI中文摘要:
许多运营问题是具有大型组合动作空间和相互依存可行性约束的约束序列决策过程。混合整数线性规划(MILP)可灵活处理此类约束,但在随机环境中扩展性较差。深度强化学习(DRL)有望实现可扩展的决策规则,但现有方法要么对约束进行惩罚而非强制实施,要么依赖的可行性机制在约束相互作用时会失效。我们通过在策略中嵌入可微凸优化模块来弥合这一差距:神经网络提出连续动作目标,二次规划将其投影到松弛可行集,双变量驱动的整数映射在保持可行性的同时恢复整数性。在给定可微模拟器的情况下,该策略使用路径梯度从采样轨迹中端到端训练,同时以与MILP相似的灵活性处理硬约束。我们证明,与精确整数投影相比,我们的可行性实施具有有界误差,且确保整个可行动作空间可被触及。我们将该方法应用于具有共享资源和物料约束的多层级生产-库存规划中。在小型实例上,我们的策略实现了低于1%的平均最优性差距;在较大网络中,其性能较最先进的层级基础库存策略高出最多9.75%,较滚动时域多阶段随机规划高出至少7.7%。在ASML的行业规模案例研究中,与已知最佳基准策略相比,它将平均成本降低了最多3.22%,且在规划难度最大的系统——即产能紧张、需求波动大的系统中,节省幅度最大。更广泛地说,我们的研究表明,在实际中普遍存在的具有相互依存硬约束的序列决策问题中,DRL可实现具有经济意义的节省。
英文摘要:
Many operational problems are constrained sequential decision processes with large, combinatorial action spaces and interdependent feasibility constraints. Mixed-integer linear programs (MILPs) handle such constraints flexibly but scale poorly in stochastic environments. Deep reinforcement learning (DRL) promises scalable decision rules, but existing methods either penalize constraints rather than enforce them, or rely on feasibility mechanisms that break down once constraints interact. We bridge this gap by embedding a differentiable convex optimization module inside the policy: a neural network proposes continuous action targets, a quadratic program projects them onto the relaxed feasible set, and a dual-informed integer mapping restores integrality while preserving feasibility. Given a differentiable simulator, the policy trains end to end from sampled trajectories using pathwise gradients, while handling hard constraints with similar flexibility to MILPs. We show that our feasibility enforcement has bounded error relative to an exact integer projection and ensures the entire feasible action space is reachable. We apply the method to multi-echelon production-inventory planning under shared resource and material constraints. Our policy attains an average optimality gap below 1% on small instances. It further outperforms state-of-the-art echelon base-stock policies by up to 9.75% and a rolling-horizon multi-stage stochastic program by at least 7.7% in larger networks. On an industry-scale case study from ASML, it reduces average cost by up to 3.22% relative to the best-known benchmark policy. The savings are largest where planning is hardest: in tightly capacitated systems with high demand variability. More broadly, our work shows that DRL can deliver economically significant savings in sequential decision problems with interdependent hard constraints, which are widespread in practice.