AI 中文总结
针对工业微电网中炼钢过程负荷调度的多级耦合可行性问题,提出嵌入过程知识的安全深度强化学习框架,实现零过程损失,购电成本较基准方法大幅降低。
AI 中文摘要
炼钢过程负荷(SPLs)是灵活资源,可提升工业微电网中本地可再生能源利用率并降低购电成本。但强多级耦合特性使得当前决策会影响后续可行性,传统深度强化学习难以在降低成本的同时保证生产全过程的可行性。本文提出一种嵌入过程知识的安全深度强化学习框架,用于工业微电网中SPL的实时调度。具体而言,构建了无损主动前沿动作空间,过程距离引导的动作处理机制会根据过程距离和演员网络的安全动作偏好重新分配被排除动作的概率;建立递归过程可行性以保证可执行性和后续可行性;此外,通过修正预算和原始对偶更新将预期过程修正距离融入PPO,将过程知识内化到原始策略中,同时推导的界定量化了原始策略对安全处理的依赖性。基于真实数据的案例研究表明,在可接受的计算时间内,相较于基于规则的调度和滚动混合整数线性规划(rolling MILP),该方法实现了零过程损失,购电成本分别降低49.2%和25.9%。
英文摘要
Steelmaking process loads (SPLs) are flexible resources that enhance local renewable-energy utilization and reduce electricity procurement costs in industrial microgrids. However, strong multistage coupling makes current decisions affect subsequent feasibility, challenging conventional deep reinforcement learning to reduce costs while maintaining process feasibility throughout production. This paper proposes a process-knowledge-embedded safe deep reinforcement learning framework for the real-time dispatch of SPLs in industrial microgrids. Specifically, a lossless active-frontier action space is constructed, and a process-distance-guided action-processing mechanism reallocates excluded-action probabilities according to process distance and the actor's safe-action preference. Recursive process feasibility is established to guarantee admissible execution and feasible continuation. Furthermore, the expected process-correction distance is incorporated into PPO through a correction budget and a primal-dual update to internalize process knowledge into the raw policy, while a derived bound quantifies the raw policy's dependence on safety processing. Case studies using real-world data demonstrate zero process losses, electricity-cost reductions of 49.2% and 25.9% relative to rule-based scheduling and rolling MILP, respectively, within an acceptable computation time.
Comments10 pages