面向多区域住宅建筑节能HVAC控制的安全深度强化学习
Safe Deep Reinforcement Learning for Energy-Efficient HVAC Control in Multi-Zone Residential Buildings
AI总结:
该研究针对多区域住宅建筑HVAC控制的能耗-安全问题,提出带安全认证的深度强化学习框架,在仿真中训练PPO与SAC智能体,验证了其在降低能耗、减少舒适度违规方面的有效性及安全保证的可行性。
AI中文摘要:
HVAC系统占建筑能耗的很大一部分,传统控制策略难以同时协调多个区域的能耗-舒适度权衡。强化学习(RL)提供自适应、数据驱动的控制方式,可随时间优化性能,但将训练得到的神经网络控制器部署到安全关键型建筑系统中仍具挑战性,原因在于缺乏形式化安全保证。本文提出一种用于多区域住宅HVAC控制的安全认证深度RL框架,在EnergyPlus/Sinergym仿真环境中训练近端策略优化(PPO)和软 Actor-Critic(SAC)智能体,以最小化能耗并维持热舒适度;基于现有神经网络Lipschitz常数计算工具,对PPO策略开展训练后安全认证,采用基于Lipschitz的前向不变性分析,确保约束满足。在八区域可变制冷剂流量(VRF)测试平台的年度仿真周期内评估两种智能体,结果显示:与基于规则的控制相比,PPO智能体实现67%的舒适度违规降低,SAC智能体实现27.6%的能耗降低;PPO策略满足形式化安全认证,余量为2.003℃,这些结果证明将强化学习与训练后安全验证结合用于多区域建筑控制的可行性。
英文摘要:
HVAC systems represent a major share of building energy consumption. Traditional control strategies are limited in coordinating energy-comfort tradeoffs across multiple zones simultaneously. Reinforcement learning (RL) offers adaptive, data-driven control that optimizes performance over time. However, deploying learned neural network controllers in safety-critical building systems remains challenging due to lack of formal safety guarantees. We propose a safety-certified deep RL framework for multi-zone residential HVAC control. Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) agents are trained in an EnergyPlus/Sinergym simulation to minimize energy consumption while maintaining thermal comfort. Post-training safety certification is performed on the PPO policy using Lipschitz-based forward invariance analysis, building on existing tools for the computation of Lipschitz constants for neural networks, to guarantee constraint satisfaction. Both agents are evaluated over an annual simulation cycle in an eight-zone variable refrigerant flow (VRF) testbed. The PPO agent achieves 67\% comfort violation reduction compared to rule-based control, while the SAC agent achieves 27.6\% energy savings. The PPO policy satisfies formal safety certification with a margin of $2.003^\circ$C. These results demonstrate the feasibility of combining reinforcement learning with post-training safety verification for multi-zone building control.