arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习调节而非循环:软 Actor-评论家算法恢复逆变器式热泵控制

Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump Control

Faizan Ahmed, Aniket Dixit, James Brusey

arXiv 2608.09453首次发表:更新:

发表机构

Centre for Computational Science and Mathematical Modelling, Coventry University(考文垂大学计算科学与数学建模中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对住宅热泵压缩机磨损问题,通过在强化学习奖励中加入压缩机磨损项,发现软Actor-评论家算法(SAC)可学习到逆变器式连续调节策略,在控制成本小幅增加的同时大幅降低热不舒适性并消除压缩机循环。

AI 中文摘要

开关循环是住宅热泵压缩机磨损的主要原因,但面向建筑的强化学习(RL)控制器通常仅优化能源成本与热舒适性,忽略学习到的策略循环程度。我们在控制奖励中加入单位压缩机磨损项,研究所得行为如何依赖于RL算法。在针对BOPTEST最佳水力热泵案例的相同马尔可夫决策过程中训练软Actor-评论家算法(SAC)和近端策略优化算法(PPO),发现SAC学习到持续调节策略,使压缩机始终处于运行状态——即逆变器驱动热泵的运行原理——实现每日零启动;而PPO则退化为开关控制,其循环次数超过基线。在BOPTEST模拟器上,SAC策略可将热不舒适性降低多达90.7%,仅伴随11.5%的成本增加,同时消除了所有基线循环。

英文摘要

On--off cycling is the main cause of compressor wear in residential heat pumps, yet reinforcement learning (RL) controllers for buildings typically optimise only energy cost and thermal comfort, ignoring how much the learned policy cycles. We add a levelised compressor-wear term to the control reward and study how the resulting behaviour depends on the RL algorithm. Training Soft Actor---Critic (SAC) and Proximal Policy Optimisation (PPO) on an identical Markov decision process for the BOPTEST bestest hydronic heat pump case, we find that SAC learns a continuous modulation policy that keeps the compressor permanently engaged---the operating principle of an inverter-driven heat pump---achieving zero start-ups per day, whereas PPO collapses to bang-bang control that cycles more than the baseline. On the BOPTEST emulator the SAC policy cuts thermal discomfort by up to 90.7% for an 11.5% cost increase, while eliminating all baseline cycling.

Comments9 pages, 4 figures, Accepted at UKCI 2026, Will be published in Springer

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑