AI 中文总结
本文提出DiffAPQP框架,实现电力系统决策导向学习的高效训练,在IEEE 118节点系统上较CvxpyLayers获得最高6.39倍的端到端训练加速,内存降低约50%且运行成本相当。
AI 中文摘要
决策导向学习(Decision-focused learning, DfL)旨在训练预测模型,使其与下游决策后果(如电力系统运行成本)对齐,但其在实际电力网络中的应用受限于训练过程中需反复求解并对大型优化问题求导的需求。本文提出DiffAPQP,这是一个基于求解器的灵活框架与开源Python包,用于结合仿射参数二次规划的可扩展DfL。为加速前向传播,DiffAPQP自动将CVXPY编写的二次电力系统模型规范化为可求导表示,并利用训练期间的重复求解结构,通过求解器热启动和求解器-数据更新实现加速。对于后向传播加速,本文建立了通过完整KKT系统求导与通过消除非活跃不等式约束得到的简化系统求导的等价性。对于仅依赖最优值的训练损失,本文进一步推导了基于包络定理的梯度,避免求解伴随KKT系统,从而缩短后向传播时间。据所知,本文是首个在IEEE 118节点系统上开展的基于求解器的端到端DfL演示,该系统包含24小时耦合的经济调度与再调度时域。在Linux机器上使用匹配的SCS和Clarabel后端时,DiffAPQP相比CvxpyLayers实现了2.27倍至3.58倍的闭环端到端DfL训练加速,以及3.62倍至4.38倍的反事实端到端DfL训练加速;最优求解器配置将这些加速比分别提升至3.91倍(每轮 epoch 从38.65分钟降至9.55分钟)和6.39倍(每轮 epoch 从10.73分钟降至1.68分钟)。此外,DiffAPQP将峰值内存使用量降低了约50%,同时保持了与CvxpyLayers相近的运行成本。
英文摘要
Decision-focused learning (DfL) trains forecasting models to align downstream decision consequences, such as power-system operating costs. However, its application to realistic power networks is limited by the need to repeatedly solve and differentiate large optimization problems during training. This paper presents DiffAPQP, a solver-flexible framework and open-source Python package for scalable DfL with affine-parametric quadratic programs. To accelerate the forward pass, DiffAPQP automatically canonicalizes quadratic power-system models written in CVXPY into a differentiation-ready representation and takes advantage of the repetitive solving structure through solver warm-start and solver-data update during training. For the backward pass acceleration, we establish the equivalence between differentiation through the full KKT system and a reduced system obtained by eliminating inactive inequality constraints. For training losses depending solely on the optimal value, we further derive an envelope-theorem-based gradient that avoids solving an adjoint KKT system, resulting in eligible backward time. To our knowledge, this work presents the first solver-based end-to-end DfL demonstration on the IEEE 118-bus system with a 24-hour coupled economic-dispatch and redispatch horizon. Under matched SCS and Clarabel backends on a Linux machine, DiffAPQP achieves $2.27\times$--$3.58\times$ closed-loop and $3.62\times$--$4.38\times$ counterfactual end-to-end DfL training speedups over CvxpyLayers. The best solver configurations increase these speedups to $3.91\times$ (from 38.65 to 9.55 min/epoch) and $6.39\times$ (from 10.73 to 1.68 min/epoch), respectively. Additionally, DiffAPQP reduces peak memory usage by approximately $50\%$, while keeping similar operating costs as CvxpyLayers.