arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11933math.OCcs.LG

迈向可持续氢能系统:基于模型预测控制与强化学习的供应链优化

Towards Sustainable Hydrogen Systems: Supply Chain Optimization with Model Predictive Control and Reinforcement Learning

Mahammad Valiyev

AI总结:

本文针对可再生能源驱动的氢能供应链,比较了基于规则、模型预测控制和强化学习等四种控制策略,发现MPC经济性最优,强化学习无需预测也能稳健运行,为运营策略选择提供指导。

AI中文摘要:

氢能供应链预计将在未来低碳能源系统中发挥核心作用,通过促进可再生能源整合、长期储能以及工业和交通部门的脱碳。然而,其运行面临可再生能源发电波动、电价波动、不确定的氢需求以及与电解槽、储能和电网交互相关的工程约束等挑战。随着氢能基础设施向商业部署扩展,运营策略必须在动态和不确定条件下平衡经济性能、可靠性和可持续性。本文研究并比较了四种用于可再生能源驱动的氢能供应链的控制方法:基于规则的控制器(RBC)、模型预测控制(MPC)、无预测的强化学习(RL-NF)以及带预测增强观测的强化学习(RL-F)。所有方法均在统一的、物理现实的框架内进行评估,该框架包含电解槽最小负荷和爬坡率约束、电池和氢储能动态、电网进口限制以及一致的经济假设,从而在相同运行条件下实现公平比较。仿真结果表明,MPC通过利用短期预测来协调储能、减少电网依赖并提高效率,实现了最高的经济性能。RL-NF在没有未来信息的情况下表现出稳健且具有竞争力的性能,凸显了基于学习的方法从经验中发现有效策略的能力。RL-F并未持续优于其无预测对应方法,这表明预测不确定性和增加的状态复杂性可能限制预测增强学习的效果。这些结果为未来氢能系统中运营控制策略的选择提供了指导。

英文摘要:

Hydrogen supply chains are expected to play a central role in future low-carbon energy systems by enabling renewable energy integration, long-duration storage, and decarbonization of industrial and transportation sectors. However, their operation is challenged by renewable generation variability, electricity price fluctuations, uncertain hydrogen demand, and engineering constraints associated with electrolyzers, energy storage, and grid interaction. As hydrogen infrastructure expands toward commercial deployment, operational strategies must balance economic performance, reliability, and sustainability under dynamic and uncertain conditions. This paper investigates and compares four control approaches for a renewable-powered hydrogen supply chain: a rule-based controller (RBC), model predictive control (MPC), reinforcement learning without forecasts (RL-NF), and reinforcement learning with forecast-augmented observations (RL-F). All methods are evaluated within a unified, physically realistic framework incorporating electrolyzer minimum-load and ramp-rate constraints, battery and hydrogen storage dynamics, grid import limits, and consistent economic assumptions, enabling a fair comparison under identical operating conditions. Simulation results show that MPC achieves the highest economic performance by exploiting short-term forecasts to coordinate storage, reduce grid dependence, and improve efficiency. RL-NF demonstrates robust and competitive performance without future information, highlighting the capability of learning-based methods to discover effective policies from experience. RL-F does not consistently outperform its no-forecast counterpart, suggesting that forecast uncertainty and increased state complexity can limit forecast-augmented learning. The results provide guidance for selecting operational control strategies in future hydrogen energy systems.

↑